Dialogue and Discourse 5(1) (2014) 1-36 doi: 10.5087/dad.2014.101 Corpus-driven Semantics of Concession: Where do Expectations Come from?∗ Livio Robaldo ROBALDO@DI.UNITO.IT Department of Computer Science University of Torino Torino, Italy Eleni Miltsakaki ELENIMI@SEAS.UPENN.EDU Institute for Research in Cognitive Science University of Pennsylvania Philadelphia, USA Editor: Raquel Fernández Abstract Concession is one of the trickiest semantic discourse relations appearing in natural language. Many have tried to sub-categorize Concession and to define formal criteria to both distinguish its subtypes as well as for distinguishing Concession from the (similar) semantic relation of Contrast. But there is still a lack of consensus among the different proposals. In this paper, we focus on those approaches, e.g. Lagerwerf (1998), Winter and Rimon (1994), and Korbayova and Webber (2007), assuming that Concession features two primary interpretations, “direct” and “indirect”. We argue that this two way classification falls short of accounting for the full range of variants identified in naturally occurring data. Our investigation of one thousand Concession tokens in the Penn Discourse Treebank (PDTB) reveals that the interpretation of concessive relations varies according to the source of expectation. Four sources of expectation are identified. Each is characterized by a different relation holding between the eventuality that raises the expectation and the eventuality describing the expectation. We report a) a reliable inter-annotator agreement on the four types of sources identified in the PDTB data, b) a significant improvement on the annotation of previous disagreements on Concession-Contrast in the PDTB and c) a novel logical account of Concession using basic constructs from Hobbs (1998)’s logic. Our proposal offers a uniform framework for the interpretation of Concession while accounting for the different sources of expectation by modifying a single predicate in the proposed formulae. Keywords: Concession, Contrast, Discourse relations, Penn Discourse Treebank 1. Introduction Our previous work on the semantic annotation of discourse relations in the Penn Discourse Tree- bank (Miltsakaki et al. (2008)) confirmed the well-known problem of achieving satisfactory inter- annotator agreement at the discourse level. As we move deeper into making semantic distinctions between eventualities participating in discourse relations, it becomes increasingly even more diffi- cult to achieve reliability, even for well-studied relations such as Concession. While the less fine ∗. Financial support to Eleni Miltsakaki by the NSF IIS-0803538 grant is gratefully acknowledged. The work of Livio Robaldo has been funded by the Ateneo-San Paolo project number TO call03 2012 0046: “The role of visual im- agery in lexical processing (RVILP)”. c©2014 Livio Robaldo and Eleni Miltsakaki Submitted 01/14; Accepted 12/14; Published online 12/14 ROBALDO AND MILTSAKAKI semantic level described as COMPARISON in the PDTB enjoyed high inter-annotator agreement, the distinction between its subtypes, i.e. Contrast and Concession, proved more challenging. A non-negligible 30% of tokens had to be adjudicated because in these instances one annotator picked Concession and the other Contrast. This was surprising given the reasonably clear semantic definition of the PDTB labels (cf Prasad et al. (2008) or Section 2 below). A careful analysis of the instances of disagreement as well as the examples studied so far in the literature revealed that a lot of linguistic variants could convey Concession and Contrast. Of course, such a high variability makes the distinction between the two classes, or between their subtypes, rather unclear, especially when it is carried out by non-expert annotators. The research that we report in this paper is motivated by our belief that data driven represen- tations can help us develop more precise formalization of semantic distinctions known to present challenges for annotators. Precise semantic distinctions guided by naturally occurring data will help us to build semantic representations, the reliability of which can be empirically tested. We summarize our research goals below. (1) a. How can the study of discourse relations as attested in naturally occurring data improve our understanding of the semantics of discourse relations? b. What kind of semantic representation will allow covering the rich range of variants conveying Concession and Contrast? To address these questions, our research methodology is guided by a hybrid theoretical and empirical approach. We develop formal semantic representations of discourse relations based on an analysis of large scale empirical data. Specifically, we analyze the semantic tagging and inter- annotator agreement of the discourse relations marked in the Penn Discourse Treebank (Prasad et al. (2008)). The Penn Discourse Treebank 2.0 is, to date, the largest annotation effort at the discourse level, including approximately 40,000 annotations of discourse connectives and their arguments, sense labels, and speaker attribution. In the PDTB, sense labels are grouped in four basic types of semantic relations: a) TEMPO- RAL, b) CONTINGENCY, c) COMPARISON, and d) EXPANSION. Each category has types and subtypes. The full hierarchy of senses used in the PDTB is illustrated in Prasad et al. (2008) and Miltsakaki et al. (2008). As mentioned above, our focus in this paper is on the distinction between Concession and Contrast, the two subtypes of the COMPARISON relation. As is natural when the body of the literature is large and coming from different disciplines, the interpretation of concessive relations has been addressed from several viewpoints. Mann and Thompson (1988)’s influential Rhetorical Structure Theory (RST) views relations from a functional perspective. The proposed interpretation includes the speaker’s intention and the effect that the re- lation is intended to achieve on the hearer. Grote et al. (1995) implement the theoretical insights from RST and other similarly minded proposals in the same vein into a real natural language gener- ation (NLG) system designed to generate concessive sentences from formal representations of the speaker’s beliefs and communicative intentions. Following Moore and Pollack (1992), we recognize the distinction between the intentional and informational levels of interpretation and find it problematic that the RST presumes a single relation between two discourse segments, thus conflating this distinction. In this spirit, our work extends prior work in sub-classifying discourse relations and developing formal representations of the identified classes. 2 CORPUS-DRIVEN SEMANTICS OF CONCESSION Our analysis revealed that concessive relations differ according to the source of expectation. Specifically, we identified four distinct sources of expectation: Causality, Implication, Correlation, and Implicature. The reliability of the proposed categories was evaluated with a study of inter-annotator agree- ment. In addition to confirming the reliability of the proposed distinctions, we evaluated the merits of this proposal over existing denial-based approaches which treat all eventualities that trigger ex- pectations uniformly. Specifically, we extracted 200 problematic PDTB tokens that had previously been marked as tokens of inter-annotator disagreement between Concession and Contrast. These tokens were re-annotated by two new annotators. The high inter-annotator agreement in this chal- lenging task provided further evidence for the validity of our proposal. Finally, we developed a formal account of Concession, grounded on our sub-classification, using basic semantic constructs from Davidson (1967) and Hobbs (1998). The resulting formulae are able to uniformly take into account the semantics of all variants of concessive statements identified in the literature. The paper is organized as follows. Section 2 gives a brief overview of prior work on Concession and Contrast, while comparing the two classes. Section 3 reviews logical accounts that have been proposed to model concessive interpretations, while Section 4 highlights some important questions that have remained open. Section 5 illustrates examples taken from PDTB that convinced us to further classify Concession into four subtypes, depending on how expectations are created. On the other hand, Section 6 reports the results of two empirical investigations carried out on PDTB instances that seems to support our analysis. In Section 7, we present, briefly, the basic semantic constructs that we use from Hobbs’s logic and outline in detail our semantic account for all but one source of expectation. The source of expectation we do not encompass in our approach is ‘Implicature’. That requires pragmatic reasoning and so it is left for future work. We conclude in Section 8. 2. Concession and Contrast Concession is a particular relation holding between the interpretation of one clausal argument that creates an expectation and another clausal argument which denies it. In English, typical discourse connectives conveying Concession are ‘but’, ‘although’, ‘however’, ‘yet’, and ‘nevertheless’. Con- cessive discourse connectives are, of course, available in other languages (König (1983)), which also have specialized words or even inflections to mark concessive relations, c.f. Dascal and Ka- triel (1977), Horn (1989), and Lagerwerf (1998). According to König (1983), this diversity in the linguistic devices used to express concession suggests that the term ‘concessive’ does not only ex- press a two-term relation, but also other possible rhetorical uses of the involved clauses. In the same spirit, Grote et al. (1995) identify three rhetorical strategies a concessive construction may be built for: convincing the hearer, preventing false implicatures, and emphasizing surprising events. We investigate the interpretation of discourse connectives only, leaving outside other linguistic or non-linguistic cues that might be used to express concession. Discourse connectives, in English and other languages, may use the same connective to express more than one type of relation. For instance, as observed in the PDTB, ‘but’ is used to express both Contrast and Concession. In line 3 ROBALDO AND MILTSAKAKI with prior work (Lakoff (1971), Spooren (1989), Grote et al. (1995), and Kehler (2002), among others), the PDTB adopts the following two definitions of Contrast and Concession1: (2) a. Contrast applies when the connective indicates that the two sentence-arguments share a predicate or property and a difference is highlighted with respect to the values assigned to the shared property. Examples are: i. [John paid $5 ] but [Mary paid $10 ]. ii. [John likes music ] but [Mary likes dancing]. iii. [I read two books ] while [you read only one ]. b. Concession applies when the connective indicates that one of the arguments describes a situation A which “creates” an expectation C, while the other asserts (or implies) ¬C, i.e. it denies the expectation. Examples are: i. Although [John studied hard ], [he did not pass the exam ]. ii. [John married Mary ] but [he loves another girl ]. iii. [We are going for a walk ] even if [it is raining ]. For now, let us indicate the relation between A and the expectation C as a kind of ‘default Implica- tion’, following Winter and Rimon (1994), and we will characterize this defeasible relation later in section 7.2. As pointed out above, Contrast and Concession are the two PDTB types of the higher level semantic category COMPARISON. In the literature, prior work has described semantic classes that would fit under ‘COMPARISON’ according to the PDTB sense tagset. For the connective “but”, specifically, there has been work analyzing its “Corrective” or “Rectification” use (Dascal and Ka- triel (1977), Lang (1984), Foolen (1991), von Klopp (1994), Winter and Rimon (1994), and others). ‘Rectification’ arises when the argument of the connective rewrites a predicate, e.g., “John is not American, but British”. In PDTB, similar cases are collapsed into the type Contrast and are not our main focus2. In this paper, we are interested primarily in Concession. The present section addresses the problem of distinguishing between Concession and Contrast. In the next section, we will address the central research question of the paper: identifying different subtypes of Concession depending on how the expectation is created. Later in the paper, we will provide some evidence that focusing on the source of expectation could also help distinguishing between Concession and Contrast. Although the definitions in 2, as well as those used in most existing schemes of discourse re- lations, appear to be rather intuitive, it is possible that the distinction is sometimes hard because it is sensitive to the context. This has been investigated and argued in the work of Lakoff (1971), Anscombre and Ducrot (1979), Lang (1984), Blakemore (1989), Winter and Rimon (1994), and Spenader and Lobanova (2009). An example, taken from Winter and Rimon (1994), is: (3) [John is quick ], but [Bill is slow]. 1. All discourse connectives annotated in the PDTB have two arguments. In the examples shown in (2), Argc is shown in boldface, Argd in italics, and the discourse connective is underlined. 2. For a more detailed discussion of contrastive interpretations see Izutsu (2008). 4 CORPUS-DRIVEN SEMANTICS OF CONCESSION According to the instructions given in (2), an annotator should tag (3) as ‘Contrast’. In fact, we may identify a common property ‘Y is-a-quality-of X’, shared by the two arguments’ meaning representations, where X is either John or Bill and Y either one of the contrastive values ‘quick’ and ‘slow’. However, as discussed by Winter and Rimon (1994), (3) may be interpreted as Concession in a context where John and Bill belong to a sport team that is known to have only quick players. In this case, the first sentence in (3) is better interpreted as an assertion A which creates the expectation C that all players, including Bill, are quick, while the second sentence in (3) explicitly asserts ¬C. As can be seen, forming a neat distinction between Contrast and Concession is even harder in real data. Spenader and Lobanova (2009) point out that the following example, marked as ‘Con- cession’ in the RST corpus (Carlson et al. (2001)), could be tagged as ‘Contrast’ by an annotator who takes the brokerage operation and Kidder as parallel elements and the profits/losses as the (symmetric) values with respect to which the parallel elements contrast. (4) [Its 1,400-member brokerage operation reported an estimated $5 million loss last year ], although [Kidder expects it to turn a profit this year]. According to Winter and Rimon (1994), the context is also responsible for the interpretation of certain concessive utterances, which, in different contexts, would be odd. For instance, (5.a) sounds odd if it is not interpreted in an appropriate context as in (5.b) (Note that we mark the argument that creates the expectation and the one that denies it as Argc and Argd respectively). (5) a. #[I love Venice ]Argc , but [I would like to be there again ]Argd . b. I always hated to visit again cities that I love. Not in the case of Venice: [I love it ]Argc , but [I would like to be there again]Argd . Context and other pragmatic considerations are not only involved in the identification of the “default Implication” that creates the expectation from the meaning of Argc. Once such Implication is identified, the concessive interpretation may be triggered in different ways. Sweetser (1990) identified three possible ways: Content (or semantic) usage, Epistemic usage, and Speech Act usage. These three classes have been recognized by many authors, among which Lagerwerf (1998) and Lang (2000). Three examples of the classes, from Lagerwerf (1998), are shown in (6). (6) a. [Connors did not use Kevlar sails ]Argd, although [he expected little wind.]Argc (Content usage) b. [Theo was not exhausted ]Argd ,although [he was gasping for air. ]Argc (Epistemic usage) c. [Mary loves you very much ]Argd , although [you already know that. ]Argc (Speech Act usage) In (6.a), Argc creates the expectation via the pragmatic “default Implication” if Connor expects little wind, then he uses Kevlar sails. Based on the observation that Theo is gasping for air, the speaker of (6.b) expects him to be exhausted. However, in (6.b), it cannot be deduced, or assumed, that if someone gasps for air, then he is exhausted. Rather, the “default Implication” has the opposite direction: if someone is exhausted, then he gasps for air. It is the fact of being exhausted that causes 5 ROBALDO AND MILTSAKAKI gasping for air, not the opposite. Lagerwerf (1998) argues that, while Content usage involves a “default Implication” which is derived deductively, Epistemic usage involves one that is derived abductively: from the observations to the causes3. Finally, in (6.c) the expectation is denied by the illocution of Argd, i.e., its speech act, rather than by its locutionary meaning. It is the fact that I tell you Argd, and not Argd itself, that is inconsistent with the expectation, created by the “default Implication” If (I know that) you already know something, I do not tell it to you. Note that it is possible to distinguish three further sub-cases of Speech Act usage: the concessive relation may involve the illocution of Argd only, the one of Argc only, or it may involve both. As argued in Winter and Rimon (1994), an occurrence of Concession holding between the speech acts of both arguments is easily obtained by using imperative mood in both: (7) [Take a chair ]Argc, but [do not sit.]Argd Lagerwerf (1998) was the first who tried to collect all data and previous analyses, e.g., Lakoff (1971) and Sweetser (1990), and propose a classification of contrastive relations in Natural Lan- guage. He distinguished three main classes: Semantic opposition, Denial of expectation, and Con- cessive opposition4. (8) a. Semantic opposition, e.g., [Greta was single] but [Prince was married.] b. Denial of expectation, e.g., examples (6.a-c). c. Concessive opposition, e.g., Shall we go to King Tsin? [King Tsin has great mu shu pork ]Argc but [China First has good dim sum.]Argd According to Lagerwerf, all classes in (8) may involve speech acts in the arguments, not only Denial of expectation (which we have seen in (6)). Semantic opposition corresponds to the PDTB definition of Contrast, e.g., (2.a) whereas both Lagerwerf’s Denial of expectation and Lagerwerf’s Concessive opposition are included in PTDB’s definition of Concession. Treating these as sub-types of Concession is an approach that has also been adopted in several other works, Winter and Rimon (1994) and Korbayova and Webber (2007). We focus on this treatment in the next section. For the present discussion, the critical distinction is between Lagerwerf’s Semantic opposition and Concessive opposition that we saw in the examples (8.a) and (8.c), respectively. With respect to (8.c), note how the preceding question “Shall we go to King Shin?” sets up a context that favors the concessive interpretation. This observation has been empirically verified by Spooren (1989) who conducted an experiment with English speakers who showed a clear tendency to consider China First as the preferred restaurant. However, as discussed in Lagerwerf (1998), an interpretation of Semantic opposition is, also, possible for (8.c) if a different question is used to set up the context: 3. Abduction is a form of non-monotonic defeasible reasoning. However, note that in case of Concession the deductive counterpart is also defeasible, due to “default Implication”. See also section 7.3. 4. In Lagerwerf (1998), (8.c) is indeed termed as ‘Concession’, not as ‘Concessive opposition’. In PDTB 2.0 and other related work Korbayova and Webber (2007), the definition of ‘Concession’ encompasses both Lagerwerf’s ‘Denial of expectation’ (6.a-c) and Lagerwerf’s ‘Concession’ (8.c). Korbayova and Webber (2007) rename Lagerwerf’s ‘Concession’ as ‘Concessive opposition’ to avoid misunderstandings. In this paper, we follow Korbayova and Webber (2007)’s terminology. 6 CORPUS-DRIVEN SEMANTICS OF CONCESSION (9) Which restaurant is better? [King Tsin has great mu shu pork ]Argc but [China First has good dim sum.]Argd Indeed, Lagerwerf (1998) argues that with the wh-question in (9), Semantic opposition is the only available interpretation; two restaurants are compared with respect to their properties. From the discussion above it becomes clear that, in some cases, the crucial distinction between a concessive versus contrastive interpretation is entirely context dependent. Different contexts, e.g., an introductory wh- versus yes/no question, enables different interpretations of the same statements, focusing on a symmetric versus asymmetric perspective with respect to the compared items. Furthermore, both Concessive and Contrastive connectives compare two clausal arguments and highlight some kind of disagreement between the facts they each assert. Contrast is a symmetric relation between two opposite facts. Concession, on the other hand, features some kind of “direc- tionality”: one of the two arguments asserts a fact that triggers an expectation and the other argument overrides it later. Since context strongly influences the identification of the asymmetric roles of the arguments, in several cases the distinction could be rather subtle. Although such instances are not pervasive, in the inter-annotator study that we report in Section 6, we observed that in some cases, both a contrastive and a concessive interpretation could be built for the same token. 3. Concessive interpretations In the previous section, we discussed prior work on Concession with a focus on its difference from Contrast. In this section, we review the literature on Concession focusing on the various interpreta- tions that have been proposed. Lagerwerf (1998) identified two types of Concession that he termed as Denial of expectation and Concessive opposition. A similar sub-categorization has also been outlined in prior work on the logical formalization of concessive relations, given in Abraham (1991), Winter and Rimon (1994), Korbayova and Webber (2007). The corresponding formalizations model a direct and a less direct relation between the triggered expectation and the content of the textual span that denies it. Korbayova and Webber (2007) explain the distinction giving the two examples in (10). For simplicity, we will refer to them as instances of ‘Direct Concession’ (10.a) and ‘Indirect Concession’ (10.b). (10) a. Although [Greta Garbo was considered the yardstick of beauty ]Argc , [she never married ]Argd . (Denial of Expectation ≡ Direct Concession) b. Although [he does not have a car ]Argc , [he has a bike ]Argd . (Concessive Opposition ≡ Indirect Concession) In (10.a), a general “default Implication” is presupposed, paraphrasable as Beautiful women usually get married. Because of this rule, Argc directly triggers the expectation that Greta Garbo got married. This expectation is explicitly denied by Argd. Example (10.b) is different. In this case, “not having a car” does not imply “not having a bike”, i.e., no defeasible rule holds between the two arguments. According to Lagerwerf (1998)’s terminology, in this case we can identify a ‘Tertium Comparationis’, i.e., a proposition entailed by Argc with its negation entailed by Argd. This proposition is presumably “he is mobile”. Thus, two general rules can be identified from the arguments to the Tertium Comparationis: not having a 7 ROBALDO AND MILTSAKAKI car implies being less mobile and having a bike implies being mobile. As discussed in the previous section, the identification of the Tertium Comparationis strongly depends on context, e.g., it could be induced by a suitable introductory question. In the same spirit, Sanders et al. (1992), Sanders et al. (1993) argue that Denial of Expecta- tion is ‘causal’, while Concessive Opposition is ‘additive’. The same conceptual difference has been sharpened in other works on discourse coherence, e.g., Asher (1993) (‘structural relation’ versus ‘non-structural relation’), Kehler (1994) (‘common topic’ versus ‘coherent situation’), and Pander Maat (1998) (‘causal relation’ versus ‘comparative relation’). How has this basic intuition on the relationship between expectation and denied expectation been formalized? Francez (1995) proposes bilogic which uses two semantic structures, the standard and the actual world. The contrast between the two worlds gives rise to what we characterize here as Concession and models the difference between Concession and Contradiction in terms of whether the statements are evaluated in different worlds. Winter and Rimon (1994) agree with Francez (1995) on the basic intuition but propose to analyze Concession as presupposition failure. They combine presupposition failure with the possibility and necessity operators of modal logic to define the semantics of what they classify as direct (‘although’, ‘even though’, ‘yet’, ‘nevertheless’) and indirect concessive connectives (‘but’)5. Winter and Rimon (1994)’s formal account of the two cases is shown in (11). In (11), p and q are the propositions denoted by Argc and Argd respectively, and � is the standard possibility operator defined in modal logic. In case of Direct Contrast (our Direct Concession), the expectation is identified by ¬q, while in case of Indirect Contrast (our Indirect Concession), a third proposition r is assumed to exist, which clearly refers to the Tertium Comparationis. Its existence is implied by q while its negation is implied by p. (11) Direct Contrast: (p ∧ q) ∧ �(p→¬ q) Indirect Contrast: (p ∧ q) ∧ ∃r[�(p→¬ r) ∧ (q → r)] Note that the ‘→’ is not the standard predicate logic operator of Implication. While the details of its characterization are too complex to review here, note that the ‘→’ denotes an underspecified “cognitive reasoning” relation and it must not be confused with the standard first order logic en- tailment. In Winter and Rimon (1994), the standard first order logic entailment is formalized via the symbol ‘⇒’. Several properties are asserted on ‘→’: it is reflexive, it holds that either ‘→’ is not transitive or that ‘⇒’ is not a special case of ‘→’, and, most importantly, ‘→’ is not defeasible. Winter and Rimon (1994) model defeasibility in terms of possible worlds, so that the consequent of ‘→’ holds when its antecedent holds, but not all cognitive implications that may be extracted from the sentence’s meaning are asserted in the current world. In order to assert that ‘p→¬ q’ (and ‘p→ ¬ r ’) only weakly hold, the implication is asserted as possible, via the modal operator ‘�’, while q (and ‘q→ r ’) are asserted as true. However, as Winter and Rimon acknowledge, some cases are problematic for such an account. For instance: (12) [John walks slowly ]Argc , but [he walks. ]Argd 5. For Winter and Rimon (1994), although, even though, yet, nevertheless have ‘restrictive’ meaning and only but is ‘non-restrictive’. This one-to-one correspondence between semantic descriptions and connectives breaks quickly when we look at empirical data. In PDTB, several connectives, including ‘but’ and other concessive connectives have more than one interpretation. The connective but, for example, has been annotated with seven sense tags. 8 CORPUS-DRIVEN SEMANTICS OF CONCESSION Suppose (12) is uttered in a context where John had a surgical operation. p may be taken as the eventuality “John walks slowly”, q as “John walks”, and ¬r as the expectation “the operation was not a success”. However, “John walks slowly” is clearly a particular case of “John walks” (p⇒q) and from this we infer that p→r holds in the current state of information. In other words, “John walks slowly” cognitively implies “the operation was a success”, which is clearly not the case. In order to handle such inconsistencies, Winter and Rimon (1994) have to propose further re- strictions on ‘→’, in terms of possible worlds. In our account, we propose a formal solution using non-monotonic (default) reasoning, instead of possible world semantics. Such a solution, which is also advocated by Winter and Rimon (1994) as a plausible alternative to their account, does not suffer from the problem exemplified in (12). Winter and Rimon’s ‘Direct Contrast’ appears to be a viable formalization of Lagerwerf’s ‘De- nial of Expectation’. However, among the three subcases of Denial of Expectation demonstrated above in (6), Winter and Rimon (1994) consider examples of Content Usage and Speech Act Usage only, while remaining silent with respect to occurrences of Epistemic Usage. Let us now turn our attention to Lagerwerf and how he formalizes his intuitions about non- Content ‘Denial of Expectation’. Lagerwerf (1998) proposes associating the sentences (6.b-c), copied here as (13.a-b) for convenience, with the formulae (14.a-b), respectively. (13) a. [Theo was not exhausted ]Argd , although [he was gasping for air. ]Argc (Epistemic usage) b. [Mary loves you very much ]Argd , although [you already know that. ]Argc (Speech Act usage) (14) a. ∀x[Gfb(x) > B(i, Exh(x))] b. ∀x[K(i, K(y, x)) > � ¬T(i, y, x)] Let us, first, consider the formula in (14.b). K(a, x) means that the agent a knows x while T(a1, a2, x) asserts that the agent a1 tells x to the agent a2. i and y are two constants referring to the speaker and the hearer respectively. Thus, T(i, y, x) means that the speaker tells x to the hearer. ‘>’ is the defeasible implication operator defined in Asher and Morreau (1991). It corresponds to Winter and Rimon’s operator ‘→’ when it is asserted within the scope of the possibility operator ‘�’. Formula (14.b) states that if someone knows something, I need not say it. This formalization directly mirrors the intuition about Speech Act Usage and we agree with that. Obviously, the parallel with Winter and Rimon’s is obtained by assuming p→¬q = ∀x[K(i, K(y, x)) > �¬T(i, y, x)]. On the other hand, in (14.a), Gfb and Exh are predicates denoting, respectively, the set of indi- viduals gasping for air and the exhausted ones, while B(i,Exh(x)) is an epistemic operator asserting that the speaker believes x to be exhausted. In the next section, we discuss the main challenges that these frameworks still need to address and we clarify our data-driven approach to meeting these challenges. 4. Challenges In the previous section, we looked at basic issues in building the semantics of concession and some of the basic logical accounts that were proposed in the literature, among which Winter and Rimon (1994), Lagerwerf (1998), and Korbayova and Webber (2007). 9 ROBALDO AND MILTSAKAKI In this section, we will look a little closer at the challenges that the concession data present to these accounts. We start with a summary of the key points of these accounts: (15) a. These approaches focus on the distinction between ‘Direct Concession’ and ‘Indirect Concession’. In ‘Direct Concession’, the expectation raised by one argument is ex- plicitly denied by the other. In ‘Indirect Concession’, a ‘Tertium Comparationis’ must be first identified, i.e., an intermediate proposition entailed by one argument, whose negation is entailed by the other. b. The expectation or the Tertium Comparationis is triggered by some kind of default Implication. A proper formalization of such a default Implication has been mostly neglected in literature. In some approaches, e.g., Sanders et al. (1993), it has been argued that it is a causal relation in case of ‘Direct Concession’ and a comparative one in case of ‘Indirect Concession’. c. As argued in Lagerwerf (1998), a logical account of Concession needs to be gen- eral enough to include several variants, featuring expectations that correspond to the Speech Acts of the arguments and/or an abductive, rather than deductive, use of the default Implication. Drawing from our observations in the instances of Concession in the PDTB, it becomes clear that any theory that recognizes all and only two concessive interpretations will fall short when accounting for real data. On the contrary, we argue that the effort to successfully characterize how expectations are created, rather than how they are denied, is critical. In other words, we recognize an underlying general principle similar to the ABC-scheme pro- posed in Grote et al. (1995), which has been the starting point of Korbayova and Webber (2007), where we simply assert that Argd is inconsistent with the raised expectation. Contrary to Korbay- ova and Webber (2007), who further develop that principle by specifying subtypes of Concession according to how the expectation is denied, we develop, instead, a deeper analysis on how the ex- pectation is created because, in so doing, we are able to characterize more accurately the “default Implication” mentioned in (15.b). Of course, we are still able to distinguish between Direct or Indirect Concession, but our ap- proach does not advocate any direct correspondence of logic formulae to Direct/Indirect Concession. It cannot be maintained that in all cases of Direct Concession the expectation is created by a causal rule. We have seen this even in simple examples such as (16.a-b). Unless we assume an ad-hoc context, it would be odd to assert that being a penguin “causes” not flying and that the fact that John will do his report “causes” the fact that he will do it at home. (16) a. [Penguins are birds ]Argc . Nevertheless [they do not fly. ]Argd b. [John will do his report ]Argc but [he will finish it at home. ]Argd It would, also, be hard to try to identify a Tertium Comparationis in all cases of Indirect Con- cession. Indeed, there are cases in which a Tertium Comparationis is not there to be identified. Lagerwerf (1998) suggests that an easy way to identify a Tertium Comparationis is by presenting the utterance as an answer to an appropriate question. But what could be the appropriate questions for utterances (17.a-b)? 10 CORPUS-DRIVEN SEMANTICS OF CONCESSION (17) a. Although [John ate a lot of pizza ]Argc , [he did not eat it all. ]Argd b. [Open the computer case ]Argc , but [do not touch the wires. ]Argd It is rather hard to interpret (17.a) as an answer to a particular question. Perhaps we may think about a context in which someone asks Is there some pizza left?, and the speaker replies with (17.a). In such a case, a Tertium Comparationis analysis seems possible for (17.a) : Argc is interpreted by the hearer as a negative comment to the prospect of eating, and Argd as a (stronger) positive one. In the case of (17.b), it is even harder to find a context which involves a Tertium Comparationis because the example may be uttered in any context in which someone gives instructions on how to open the computer case. To account for all the observed instances of Concession, including (17.a-b), we need a more general definition of Indirect Concession. In our account, all cases in which Argc is insufficient or irrelevant with respect to the satisfaction of speaker’s intentions are classified under the term ‘Concession-Implicature’. In (17.a), Argc is irrelevant with respect to the satisfaction of speaker’s intentions, i.e., communicating to the hearer that there is some pizza left, and could lead the latter to conclude that there is nothing for him to eat. Analogously, in (17.b), the command in Argc is insufficient, and could lead the hearer to take wrong actions. The speaker adds then further specifications by uttering Argd. On the other hand, in (8.c), Argc is interpreted by the hearer as a preference for King Tsin over China First, which does not meet the speaker’s intentions. However, the fact that the latter has another preference, and so a Tertium Comparationis may be identified, in our view is simply a special instance. If the sentence was modified to “King Tsin has great mu shu pork, but I do not want to talk about that”, the Tertium Comparationis would be less easy to identify. Finally, we agree with Lagerwerf (1998) that a proper logical account of Concession must take into account Speech Acts and abductive use of the relation triggering the expectations, but we find his formalization of the Epistemic Usage problematic for two reasons. Consider again the example in (13), repeated in (18) for convenience. (18) a. [Theo was not exhausted ]Argd ,although [he was gasping for air. ]Argc b. ∀x[Gfb(x) > B(i, Exh(x))] In (18.b), Gfb and Exh are predicates denoting, respectively, the set of individuals gasping for air and the exhausted ones, while B(i,Exh(x)) is an epistemic operator asserting that the speaker believes x to be exhausted. However, it is somehow odd to assert that the speaker believes someone to be exhausted given that he is gasping for air. The defeasible rule in (18.a) is general and therefore does not apply specifically to any particular speaker. Therefore, in the formalization, i should be most properly substituted by universal quantification over all possible believers. This is in line with Spooren (1989), Sanders (1994) and Pander Maat (1998), who identify different subjectivities that may be ascribed to the statements in a discourse. In particular, Pander Maat (1998) conducts a corpus anal- ysis showing that three perspectives of subjectivity must be distinguished: ‘objective perspective’ (the statements are objective facts that are taken to be acceptable by the speaker), ‘speaker perspec- tive’ (the belief of the statement is ascribed to the speaker), and ‘other perspective’ (the belief of the statement is ascribed to persons other than the speaker). Pander Maat (1998) proposes a revision 11 ROBALDO AND MILTSAKAKI of the hierarchy of discourse relations provided by Sanders et al. (1992) that includes a new feature specifying the perspective configuration. Secondly, the formula in (18.b) does not adhere to Lagerwerf’s intuition (Lagerwerf (1998), pp.41-42) that Epistemic Usage of Concession involves a defeasible rule which applies abductively, i.e., getting from observations to causes. In fact, the rule does not appear in his formulae, e.g., (18.b). Furthermore, the “default Implication” is also defeasible in case when it is used deductively, as in Content Usage. Therefore, if it should be asserted that the speaker believes the pre-conditions when he observes the effects, it should also be asserted that he believes the effects when he observes the pre-conditions. In our view, Lagerwerf (1998)’s valid intuition must be formalized exactly as it is stated: the formula must include an explicit defeasible rule corresponding to “being exhausted causes gasping for air”. Separately, the formula asserts that the rule yields the expectation abductively. In other words, in the formulae the assertion of the defeasible rule must remain orthogonal to its usage. To sum up, in this and the previous sections we have attempted to give a comprehensive review of the literature on Concession and the challenges that any logical account of Concession will need to address. One could view these challenges as a purely theoretical exercise in semantic theory and continue to work on them on a theoretical basis. One might argue that these theoretical challenges should be mostly irrelevant to human annotators whose task is to identify and annotate Concession in naturally occurring data. While, indeed, it may not be surprising that there are theoretical chal- lenges to be addressed in the formal treatment of concession, we were intrigued by the fact that what, for a human, seemed to be a fairly straightforward definition of Concession (Denial of ex- pectation) yielded surprisingly high disagreement among annotators. Indeed, almost 30% of PDTB tokens annotated as Concession by one annotator were annotated as Contrast by the other. Close investigation of the data helped us realize that viewing the relation with a focus on the denial of expectation made it hard for the annotators to discern the triggers of expectation. Analyzing the sources of expectation, the different types of relation that trigger them and the inferences that they allow helped them identify the relations with much improved consistency. Our work bridges the gap between corpus data and logic and our methodological approach is doing so by starting at the bottom. We started by looking at discourse connectives in the PDTB, and then built up more abstract models for deriving appropriate inferences. In the next section, we will look at Concession in the PDTB and present a data-driven analysis of the different inferences that are triggered from the range of the sources of expectations attested in the annotations of Concession. 5. Concession in the PDTB: Where do expectations come from? The PDTB corpus contains 1193 annotated instances of Concession associated with an explicit discourse connective. In order to identify the possible defeasible relations involved in concessive relations in real data, 1000 of these 1193 instances have been analyzed. Table 1 shows the distribution of all the tokens in the PDTB that were labelled6 as Concession (or any of its subtypes7) and Contrast (or any of its subtypes). There were a total of 1193 tokens 6. A full description of the sense tags used in the PDTB is given in Prasad et al. (2008) and Miltsakaki et al. (2008). 7. The PDTB distinguishes two subtypes of Concession: “expectation” and “contra-expectation”. When the clausal argument that syntactically bounds the discourse connective creates the expectation, the PDTB instance is labelled as “expectation”. Otherwise, it is labelled as “contra-expectation”. Note that this sub-categorization is orthogonal to 12 CORPUS-DRIVEN SEMANTICS OF CONCESSION labelled as Concession. The most common concessive connective is ‘but’ with 508 tokens (42% of all concessive labels), followed by ‘although’ with 154 tokens (13% of all concessive labels). The connective ‘but’ is, also, very common in relations marked as ‘Contrast’, which may have contributed to the confusion between ‘Concession’ and ‘Contrast’ that we noted earlier. Connective Concession although 154 (13%) but 508 (42.5%) even if 35 (3%) even though 72 (6%) however 77 (6.5%) nevertheless 19 (1.5%) nonetheless 17 (1.5%) still 82 (7%) though 84 (7%) while 83 (7%) yet 32 (2.5%) other 30 (2.5%) Total 1193 Connective Contrast although 157 (4.06%) but 2422 (62.8%) by contrast 27 (0.7%) even though 21 (0.51%) however 355 (9.3%) meanwhile 37 (0.95%) on the other hand 35 (0.9%) still 96 (2.45%) though 131 (3.42%) while 427 (11.16%) yet 53 (1.32%) other 76 (2.43%) Total 3856 Table 1: Concession and Contrast labels in PDTB 2.0 It seems surprising that Concession, despite having a straightforward definition, was so fre- quently confused with Contrast. For all the tokens that were annotated as either Concession or Contrast by one of the PDTB annotators, there was almost 30% disagreement. In most cases of disagreement, at least one annotator would choose Contrast over Concession because they would prefer to construct a contrastive interpretation between the created expectation and its denial, failing to see that the involved predicates were not symmetric. This section reports several PDTB instances tagged as Concession, out of 1000 selected ones, and shows that it is possible to identify four types of semantic relation that give rise to the asymmetry characterizing Concession: Causality, non-monotonic Implication, Correlation, and Implicature. The next four subsections discuss each category with corpus examples and outline their mean- ings. Support for the proposed classification of the identified sources of expectations is given by the results of an inter-annotator study that we conducted asking the annotators to label the data with the new categories (cf. section 6). 5.1 Causality In Sweetser (1990), Sanders et al. (1992), Lagerwerf (1998), and also in Prasad et al. (2008), it is assumed that in all cases of Concession the expectation comes from a (defeasible) causal rule. The next subsections argue that this is true in most, but not all, cases. An example, taken from the PDTB, in which the expectation is created by a defeasible causal rule is shown in (19). the one addressed in this paper, i.e. “direct” and “indirect” Concession. The latter concerns the way the expectation is denied, regardless of the clausal argument that creates it. 13 ROBALDO AND MILTSAKAKI (19) Although [they represent only 2% of the population]Argc , [they control nearly one-third of discre- tionary income ]Argd . In (19), Argc asserts that “they” represent a very low percentage of population. That creates the expectation that they control a (proportionally) low percentage of income. Where does this expectation come from? The obvious answer is that our world knowledge includes a general causal rule representing that a low percentage of population causes control of a small amount of income, which instantiates on Argc and creates the expectation. This causal rule is, however, defeasible, i.e., its effect may be falsified or canceled, as it is done by Argd in (19). More cases of Concession triggered by a causal rule are shown in (20). (20) a. Although [imports account for less than 1% of beer sales in Japan]Argc , [Asahi Breweries Ltd., which has been gaining share with its popular dry beer, plans to fend off Japanese com- petitors by pouring $1.06 billion into facilities to brew 50% more beer]Argd . (Causality) b. [This meeting “put in motion” procedural steps that would speed up both of these functions]Argc . But [ no specific decisions were taken on either matter]Argd . (Causality) c. [...that hung over parts of the factory]Argd even though [exhaust fans ventilated the area]Argc . (Causality) d. [A Sanwa Bank spokesman denied that the finance ministry played any part in the bank’s decision]Argc . Still [Mr. Utsumi may have a hard time convincing market analysts who have rightly or wrongly believed that the ministry played a role in orchestrating recent moves by Japanese banks]Argd . (Causality) e. [An undistinguished college student, who dabbled in zoology until he concluded that he couldn’t stand cutting up frogs, Mr. Corry wanted to work for a big company “that could do big things”]Argc But [after joining the tax department of a USX subsidiary 30 years ago, he set the modest goal of becoming tax manager by the age of 46.]Argd (Causality) In (20.a), the low percentage of sales in Japan should cause Asahi Breweries Ltd. to invest somewhere else. Similarly, in (20.b), “the procedural steps triggered by the meeting” (defeasibly) causes “taking important decisions in both of these functions” and, in (20.c), the fans should blow away whatever it was that hung over it. In (20.d), the declarations of the Sanwa Bank spokesman create the expectation that Mr. Utsumi may be on the safe side. Finally, (20.e) is particularly interesting in that the expectation is created by a conjunction of two different causes. The fact that Mr. Corry was an undistinguished college student and the fact that he had the intention to work for a big company defeasibly cause the fact that he got smart professional results. 5.2 Implication The previous subsection presented some examples of Concession where the expectations are created via abstract (defeasible) causal relations that instantiate on Argc. And, it has been pointed out that many past proposals assume that the expectation is always triggered by a causal rule. The data in the PDTB reveal that not all occurrences of Concession involve causality. An inter- esting instance is shown in (21): (21) [The prime minister,]Argd [whose hair is thinning and gray and whose face has a perpetual pallor,]Argc nonetheless [continues to display an energy, a precision of thought and a willingness to say pub- licly what most other Asian leaders dare say only privately]Argd . 14 CORPUS-DRIVEN SEMANTICS OF CONCESSION Argc describes two properties featured by the prime minister, which do not appear to cause the negation of Argd, or a Tertium Comparationis related to it. Rather, the description recalls in our minds some kind of prototypical old and tired man, of which the prime minister would be an instan- tiation. The expectation stems from the fact that the prime minister inherits all typical properties of such a prototype, among which the one of having a lazy and indolent attitude. Default inheritance from a prototype is clearly a defeasible implication, i.e., its consequent may be overridden as it is done by Argd in (21). These considerations are well-known by researchers working on Default Logics. Consider the typical example shown in (22). (22) [Penguins are birds.]Argc . Nevertheless [they do not fly]Argd . In (22), Argc suggests that penguins have the property of flying, which they inherit from the prototype of ‘bird’. This expectation is explicitly denied by Argd. In many PDTB occurrences Argc evokes a kind of prototype of which some properties are overridden by Argd. Some are reported in (23): (23) a. [So far, all the studies have concluded that RU-486 is safe.]Argc . But [“safe” in the definition of Marie Bass of the Reproductive Health Technologies Project, means “there’s been no evidence so far of mortality”]Argd . (Implication) b. [David is a pragmatist.]Argc . But [Mr. Dinkins’s sense of pragmatism often comes across more as an insider’s determination not to upset the political apple cart]Argd . (Implication) c. Although [working for U.S. intelligence]Argc , [Mr. Noriega was hardly helping the U.S. exclu- sively ]Argd . (Implication) d. Although [insider trading has long been criminal]Argc , [it has never been statutorily defined ]Argd . (Implication) e. [You can do all this ]Argd even if [ you’re not a reporter or a researcher or a scholar or a member of Congress]Argc . (Implication) In (23.a-b) it is easy to see the inheritance by default that creates the expectation. The concepts of “safe” and “pragmatic” respectively used in the sentences are not exactly the ones that are stan- darly assumed, i.e., the prototypes. Argd specifies the prominent differences with respect to such a prototype, i.e., what properties are overridden. In many cases, the prototype from which the canceled expectations are inherited is not so easy to identify. In those cases rather than thinking in terms of “inherited properties”, it is more convenient to think in terms of “necessary conditions” to which the prototype must adhere. For instance, in (23.c), working for U.S. exclusively is perceived as a necessary condition for working for U.S. intelligence. In other words, by reading Argd in (23.c) we perceive that Mr. Noriega is arguably breaking some kind of rule required by his role. Similarly, in (23.d), it seems that, in order to claim that “something is criminal”, it is necessary that “it is defined as such by the law”. Finally, in the context of (23.d) it defeasibly holds that whoever can do all this must be either a reporter, or a scholar, or a researcher, etc. 15 ROBALDO AND MILTSAKAKI 5.3 Correlation The annotators of the empirical study presented below in section 6 chose ‘causality’ or ‘defeasible entailment’ for about 70% of the occurrences of Concession taken from the PDTB. The remaining cases seem to involve different relations. Consider for instance (24): (24) [The Treasury will raise 10 billion in fresh cash by selling 30 billion of securities . . . ]Argc . But [rather than sell new 30-year bonds, the Treasury will issue 10 billion of 29 year, nine-month bonds]Argd . In (24), it does not seem that there is a general causal rule at stake. The fact that the Treasury will raise money cannot be the cause of the way it will actually do it. Arguing for the existence of a prototype evoked by Argc also seems hard, though more compatible than the causal interpretation (cf. next subsection). It seems that in examples such as (24) the expectation is created on the basis of the history of the previous similar situations. In the context, it is assumed that there are two events that usually correlate. Argc describes one of the two, and we expect the other one to co-occur based on the fact that in several similar previous situations they did so. Accordingly, the third source of expectation has been termed as ‘Correlation’. Archetypal cases of Correlation are all examples where Argc describes a kind of trend and Argd an eventuality that diverges from that trend. (25) shows some examples taken from the PDTB. (25) a. Although [the notes held at a price of 92 to 93 immediately after the reset,]Argc , [they started falling soon afterward ]Argd . (Correlation) b. [Sales of the heart drug TPA were $43.6 million, better than last year’s depressed third period when the company sold just $29.1 million of the drug ]Argc . But [TPA sales fell below levels for this year’s first and second quarter sales of $48 million ]Argd . (Correlation) c. [The LDP won by a landslide in the last election, in July 1986 ]Argc . But [less than two years later, the LDP started to crumble, and dissent rose to unprecedented heights ]Argd . (Correlation) A variant of this pattern encompasses occurrences where Argd describes an eventuality that sounds “surprising” together with the one described by Argc (cf. König (1983)). (26) shows some of such instances. In (26.a), it is “surprising” that Wedtech got rolling so late, given its start date. Similarly, in (26.b) and (26.c), it is surprising that Mr. Collor remains ‘the favorite’ and that the Journal did not mention the Reserve Fund and the creators of the money-fund concept. (26) a. Although [started in 1965 ]Argc , [Wedtech didn’t really get rolling until 1975 ]Argd . (Correlation) b. [The favorite remains Fernando Collor de Mello, a 40-year-old former governor of the state of Alagoas ]Argc . But [after building up a commanding lead, the moderate to conservative Mr. Collor has slipped to about 30% in the polls from a high of about 43% only a few weeks ago ]Argd . (Correlation) c. [Actually, about two years ago, the Journal listed the creation of the money fund as one of the 10 most significant events in the world of finance in the 20th century ]Argc . But [the Reserve Fund, America’s first money fund, was not named, nor were the creators of the money-fund concept, Harry Brown and myself ]Argd . (Correlation) 16 CORPUS-DRIVEN SEMANTICS OF CONCESSION Is Correlation a source of expectation that is really distinct from the others? As argued above, Concession may stem from Causality or Implication if the context includes a general causal or entailment rule that creates the expectation. It may then be observed that, in those cases, the event that triggers the expectation and the event that describes the expectation co-occur. Let us look at the following simpler example of Correlation: (27) [John will finish his report ]Argc , but [he’ll do it at home ]Argd . From (27), we infer that John usually does not finish his reports at home, and the present occa- sion constitutes an exception to this general trend. But it may be argued that there is a particular (unknown) reason why John never does his reports at home. Maybe his home is too noisy or the reports must be returned by the end of the work day. These reasons might cause the fact that John does not finish his reports at home. Similarly, in (24) we may think of a “prototypical Treasury” that always raises money in the same way, namely by selling new 30-year bonds. Although such considerations might indicate that Causality and non-monotonic Implication of- ten entail Correlation, in our view they should be kept distinct for two reasons. First, precisely because we do not know if there is a particular hidden reason why John does his reports at the of- fice, we should not assert its existence, unless we believe that this is the inference that the reader draws from the text, which is clearly not the case. Secondly, it has been attested beyond doubt that there are instances of concession for which no causal rule or defeasible entailment can be con- strued. There are, also, examples involving a causal rule, for which it cannot be asserted that the cause co-occurs with the effect. (28.a-b) from Winter and Rimon (1994) and Grice (1961) are cases in point: (28) a. [Take a chair ]Argc, but [do not sit.]Argd (Correlation) b. [She is poor ]Argc but [she is honest. ]Argd (Causality) In (28.a), we cannot infer a causal rule or prototype stating that encouraging someone to take a chair “causes” or “entails” an invitation to sit on it. Maybe Correlation could best model instances of concession involved in Speech Act Usage but we have not conducted a study for speech acts specifically to support any claims. Conversely, in (28.b) the expectation is created via a causal rule: poverty may be the cause driving people to criminal activity such as stealing. But, it would be wrong to infer from that causal relation that poor people tend to be dishonest, i.e., a Correlation relation. 5.4 Implicature There are cases of Concession in which the expectation is created by the pragmatics of the conver- sation. As mentioned earlier, while all occurrences belonging to this class express Indirect Contrast, not all of them involve a Tertium Comparationis. For this reason, we associate this class with a broader definition. Concession is triggered via Implicature whenever Argc is insufficient or irrelevant to the speaker’s intention. It could lead the hearer to draw unintended inferences. It seems that in such cases Argc violates a Gricean Maxim, Grice (1975). Argd adds to Argc the relevant information that the speaker wants to convey. The examples discussed in the Introduction are repeated in (29). 17 ROBALDO AND MILTSAKAKI (29) a. A: Shall we take this room? B: [It has a beautiful view ]Argc but [it is very expensive. ]Argd b. Although [John ate a lot of pizza ]Argc , [he did not eat it all. ]Argd c. [Open the computer case ]Argc , but [do not touch the wires. ]Argd In (29.a), Argc could be interpreted by the hearer as “ I (the speaker) want to rent this room”, which is not the speaker’s intention. In (29.b), Argc is irrelevant with respect to the satisfaction of speaker’s intentions, i.e., communicating to the hearer that there is some pizza left, who in the context might be looking for something to eat. Similarly, in (17.c), the command in Argc could lead the hearer into thinking that the permission to which opening the computer case extends is unconstrained. Below are some examples of Concession via Implicature taken from the PDTB: (30) a. Although [it is not the first company to produce the thinner drives ]Argc , [it is the first with an 80-megabyte drive ]Argd . b. [Also, Exxon went down 3/8 to 45 3/4 and Allied-Signal lost 7/8 to 35 1/8 ]Argc , even though [the companies’ results for the quarter were in line with forecasts ]Argd . In (30.a), Argc does not create any expectation that is inconsistent with Argd. Argd, simply, it conveys an achievement that is worth noticing in this context. Similarly, in (30.c) Argc reports some data about the stock value trend of Exxon and Allied- Signal. Argd simply stops the potential inference that their results, which are indeed independent from the stock value, were not in line with the forecasts. The PDTB does not include enough Implicature examples for analysis, so clearly more work is needed before a satisfactory treatment of this category can be offered. 6. Studies of inter-annotator agreement In this section we report two inter-annotator agreement studies that we conducted to evaluate a) the reliability of distinguishing four sources of expectations in the semantic description of Conces- sion and b) the impact of the new analysis of concession on the, previously low, inter-annotator agreement between Contrast and Concession in the PDTB. It must be pointed out that these annotation experiments ought to be considered only as prelim- inary studies of our claims, i.e., Concession is more characterized by how expectations are created rather than by how they are denied. On the other hand, in order to obtain reliable annotations we will need precise guidelines with linguistic examples and subsequent adjudication steps as suggested by Versley and Gastel (2013). Since annotating discourse relations is a rather difficult task, Versley and Gastel (2013) propose a set of linguistic tests that annotators should use in order to tag difficult non-archetypal cases. For such cases, annotators are required to perform paraphrases of the utter- ance, insertion/substitution operations of either the connective or its argument, etc. and check which aspects of the overall meaning are changed and which are not. The check should make annotators able to select the proper sense label. With respect to the ambiguity between Contrast/Concession, Versley and Gastel (2013) propose linguistic tests aiming at testing the symmetric/asymmetric role of the arguments (cf. Versley and Gastel (2013), section 4.1). 18 CORPUS-DRIVEN SEMANTICS OF CONCESSION Furthermore, since the quality of the annotations obviously does not only depend on the clarity of the guidelines, but also on how the annotators are able to apply them, Versley and Gastel (2013) suggest using a set of quantitative tests to subsequently inter-adjudicate the annotations. This is particularly strategic for discourse relations, for which all annotation schemes proposed so far in the literature appear to be intuitive with respect to sample cases, but it is not so when applied to real data, due to the strong context-sensitivity of discourse connectives (cf. (2) above). In the same spirit, Spenader and Lobanova (2009) uses χ2 to check statistically significant correlations between lexical markers and their senses. This and similar methods could be used for “filtering” discourse markers that are intuitively associated with certain senses but that, empirically, are not. For instance, Spenader and Lobanova (2009) found out that “however”, standardly taken to be a marker of Contrast, is indeed equally used in Cause-Effect relations. Nevertheless, the creation of such a reliable corpus is beyond the goal of the present paper, and it will deserve a new separate paper. The key point of our paper, we stress again, is to provide a logical formalization of concessive relations alternative to the ones proposed by Winter and Rimon (1994), Lagerwerf (1998), Korbayova and Webber (2007), and others. These proposals are essen- tially grounded on the analysis of sample sentences while our formalization is mainly guided by an empirical analysis of real data stored in the PDTB. 6.1 Annotation of expectation sources We conducted an empirical analysis on 1000 PDTB tokens of explicit connectives annotated as ‘Concession’. Two trained annotators, one of the authors and a post-doctoral researcher in linguis- tics, tagged each token with one of the four sources of Concession identified above. The post- doctoral researcher received a short tutorial about the different sources of expectation as explained and had the option to use ‘other’ if none of the suggested labels were appropriate. The option ‘other’ was not used by either annotator. The most common source of expectation comes from causal relations (41.6%), followed by Implication (28.7%), Correlation (19.4%) and Implicature (10.3%). Source Although But Total Causality 65 248 416 (41,6%) Implication 45 125 287 (28,7%) Correlation 31 87 194 (19,4%) Implicature 13 48 103 (10,3%) Table 2: Distribution of the four sources of Concession. The kappa statistic for inter-annotator agreement yielded 0.8 agreement, indicating that the de- fined categories are reliable8. In the formula below, Pr(a) is the percentage of agreement (85% of 8. Since the mid-1990s, when we saw an increased interest in producing semantic and discourse level annotations to linguistic corpora, it has been widely recognized that the highly subjective nature of semantic and pragmatic interpre- tations could yield unreliable annotations. When two annotators disagree, either one of the two annotators is wrong or the annotation schema, often a set of tag categories, is not capturing a reliable characterization. Semantic and discourse annotation efforts are renowned for the struggle to identify reliable categories that would minimize inter- annotator disagreement (among others, Carletta (1996), Di Eugenio (2000), and Poesio and Artstein (2008)). For example, in the development of the RST corpus, Carlson et al. (2001) used professional language analysts with prior 19 ROBALDO AND MILTSAKAKI the 1000 cases considered) while the percentage of each tag, i.e., Pr(e), is equal to 25%, as there are four possible sources of Concession. κ = Pr(a)−Pr(e) 1−Pr(e) = 0.85−0.25 1−0.25 = 0.8 After the computation of inter-annotator agreement, there was a brief adjudication effort that resulted in resolving any disagreements so we could compute the distribution of labels. In the cases of disagreement, we did not observe any interesting pattern to report. Table 2 shows the distribution of the four labels for the most common connectives conveying Concession, i.e., ‘But’ and ‘Although’. 6.2 Annotation of Concession vs Contrast Making a reliable distinction between Contrast and Concession was the most challenging annotation task in the PDTB, exhibiting relatively low inter-annotator agreement. In a total of 4319 instances of explicit connectives that were annotated as either Concession or Contrast, there was agreement in 3057 cases, i.e., 70.8%. For that reason, we conducted a second inter-annotator agreement study focusing only on the PDTB tokens that were annotated as either Concession or Contrast. Specifically, we extracted 200 tokens of disagreement, i.e., tokens that one annotator had labelled as Concession and the other as Contrast. We trained two annotators, not the authors, to perform the task. Both annotators were lin- guistics students who attended a two-hour seminar on the distinctions between the different sources of expectation. The definition of Contrast remained the same as in the original PDTB annotation. For each token, they were instructed to choose one of four annotation labels: a) Contrast, b) Con- cession, c) COMPARATIVE, and d) Other. They were allowed to use the label COMPARATIVE when they could not decide between Contrast and Concession and Other when they thought that the example belonged to a different semantic class. The results are reported in Table 3. The distributions of the two annotators are almost equal. But, of course, this does not mean that we obtained almost 100% agreement. Indeed, there are only six instances that have been labelled as ‘Contrast’ by both annotators. For all other instances, either both annotators chose ‘Concession’ or they assigned different labels. Each annotator used the label ‘Other’ a single time, but not for the same instance. The label ‘COMPARATIVE’ has never been selected. Therefore, annotators agreed on 161 tokens (80.5%), most of which (155 tokens) have been labelled by both annotators as ‘Concession’. The kappa score is: k = Pr(a)−Pr(e) 1−Pr(e) = 0.805−0.25 1−0.25 = 0.74 However, this kappa cannot be compared with the statistics of the original PDTB annotators, as they had more labels to choose from when they performed the annotations. But this is not critical for our experience in data annotation and they only achieved kappa 0.60 for annotating RST-style discourse relations (in- cluding concession) reaching maximum kappa 0.75 after the annotators had worked together for a week. Explaining to annotators what to do does not guarantee agreement even if they are trained. Inter-annotator studies are, therefore, crucial for the evaluation of the reliability of the suggested semantic categories. 20 CORPUS-DRIVEN SEMANTICS OF CONCESSION Label Student1 Student2 Contrast 25 24 Concession 174 175 COMPARATIVE 0 0 Other 1 1 Table 3: Distribution of Concession vs Contrast. purposes because we are only interested in evaluating the possible gain of analyzing Concession in terms of the four sources of expectation on the annotation of Contrast versus Concession. The strong preference towards Concession was indeed expected. We recall that the 200 instances were selected among those that were ambiguous between Concession and Contrast in the original PDTB annotation. Intuitively, it is somehow unlikely that such doubtful cases were conveying a symmetrical relation, which should be rather easy to identify. In other words, it is possible that the PDTB annotators could not reliably identify the underlying relations that gave rise to expectations. Focusing on the types of relations that give rise to expectations made it clearer that unlike Con- cession, an asymmetrical relation, Contrast involves a symmetrical relation between a common shared predicate receiving different values. Concession, on the other hand, is always an asymmet- rical relation, which relies on understanding the underlying relation of two events, not mentioned explicitly. Understanding the nature of the relation that gives rise to an expectation (causality, im- plication, correlation, implicature) highlights the asymmetry inherent in Concession. Therefore, in the same spirit as Versley and Gastel (2013), section 4.1, who propose linguistic tests aiming at testing the symmetric/asymmetric role of the arguments for disambiguating between Contrast and Concession, focusing on the sources of expectation could be perhaps taken as a se- mantic/pragmatic test for the very same task. Looking at the instances of persisting disagreement, we observed that in most of these cases, it was possible to construct both a contrastive and a concessive interpretation. Consider for example, token (31). In this case, both a concessive and a contrastive interpretation can be built. A contrastive interpretation can be built by juxtaposing the predicates aware but not responsible. A concessive interpretation can be built if the reader assumes that the President Waldheim knew about the killings before they happened and did nothing to prevent them. In this context, asserting that he was not responsible for the killings creates the expectation that he was not aware of them. (31) [London has concluded that Austrian President Waldheim wasn’t responsible for the execution of six British commandos in World WarII ]Argc , although [he probably was aware of the slayings]Argd . It is possible that in some cases, better understanding of the context might help in disambiguat- ing the intention of the author. On the other hand, it is also possible that both interpretations are entertained by the reader. Since we did not give the annotator the option to annotate with a double tag Concession-Contrast, we do not know if such a tag would be used. While further studies would be required to evaluate the impact of the proposed analysis on a big- ger scale, these results offer strong support in favor of looking closer at sources of expectation when analyzing Concession. With these encouraging results, we set out to develop a logical account that 21 ROBALDO AND MILTSAKAKI would most elegantly capture the semantics of Concession while making appropriate distinctions for the identified (semantic) sources of expectation. 7. Semantics of Concession This section proposes a logical account for the occurrences of Concession in which the source of the expectation is either Causality, non-monotonic Implication, or Correlation. A proper formal treatment of Concession via Implicature is seen as the object of future work. Rather than designing new ad-hoc logical constructs to handle the semantics of Concession, we make the effort to formalize our insights using an existing logical framework, if possible. The framework that allowed us to give the most elegant account has been defined in Hobbs (1998) and several other earlier publications by the same author9. Hobbs defines a wide-coverage logic for Natural Language semantics based on the notion of reification Davidson (1967), Bach (1981). It implements a fairly large set of linguistic and semantic concepts including sets, composite entities, scales, change, causality, time, event structure, etc., into an integrated first order logical formalism. Hobbs’ framework includes all ingredients needed to properly represent the concepts introduced in the previous sections, in particular the possibility of defining defeasible relations. In addition, Hobbs’ modular logic can be used to study the semantics of the connectives independently of the semantics of the arguments Argc and Argd. Finally, we also show that Hobbs’ framework is a suitable choice for the easy integration of other insights offered in the literature, such as Lagerwerf (1998)’s, and the extension to the semantics of other discourse connectives. Interestingly, Lagerwerf (1998)’s work as well as several other researchers’ work on discourse semantics, is based on the taxonomy of coherence relations proposed by Sanders et al. (1992), which is in turn based on Hobbs’ notion of “discourse coherence” Hobbs (1991). The following subsection briefly describes Hobbs’ logical framework, with a particular focus on the ingredients needed to handle concessive relations. Our proposal for a logical account of Concession will be illustrated in 7.2. 7.1 Hobbs’ logical framework Hobbs (1998) proposed a wide coverage logical framework for NL semantics centered on the notion of Reification. Reification allows a wide variety of complex natural language (NL) statements to be expressed in Predicate Logic. NL statements are formalized such that events, states, etc., correspond to constants or quantifiable variables of the logic. In other words, the states and events denoted by these constants as well as the variables are things in the world. Hobbs uses the term ‘eventuality’ to denote the reification of both a state or an event. Hobbs distinguishes two parallel sets of predicates: primed and unprimed. The unprimed pred- icates are standard first order predicates commonly used in logical representations. For example, (give a b c) asserts that a gives b to c in the real world. The primed predicate represents the reifica- tion of the corresponding unprimed relation. The expression (give′ e a b c) says that e is a giving event by a of b to c. Eventualities may be possible or actual. In Hobbs, this distinction is repre- sented via a unary predicate Rexist that holds for eventualities really existing in the world. To give 9. See http://www.isi.edu/∼hobbs/csknowledge-references/csknowledge-references.html and http://www.isi.edu/∼hobbs/csk.html. 22 CORPUS-DRIVEN SEMANTICS OF CONCESSION an example cited in Hobbs, if I want to fly, my wanting really exists, but my flying does not. This is represented as: (Rexist e) ∧ (want′ e I e1) ∧ (fly′ e1 I) Eventualities can be treated as the objects of human thoughts. Reified eventualities are in- serted as parameters of such predicates as believe, think, want, etc. Reification can be applied recursively. The fact that John believes that Jack wants to eat an ice cream is represented as an eventuality e such that it holds10: (Rexist e) ∧ (believe′ e John e1) ∧ (want′ e1 Jack e2) ∧ (eat′ e2 Jack Ic) ∧ (iceCream′ e3 Ic) Hobbs’ logic distinguishes between specific eventualities, like “Fido is barking”, and general or abstract types of eventualities, like “Dogs bark”. They are not treated as radically different kinds of entities. At some level, they are both eventualities that can be the content of thoughts. To this end, the logical framework includes the notion of typical element (from Hobbs (1995) and Hobbs (1998)). The typical element of a set is the reification of the universally quantified variable ranging over the elements of the set (cf. McCarthy (2002)). Typical elements are first-order individuals. Their introduction is motivated by the need of moving from the standard set theoretic notation in Predicate Logic: (forall (x) (iff (member x s) (p x))) to a simple statement that p is true of a “typical element” of s. In Hobbs’ notation, the typical element t of a set s satisfies the predicate (typelt t s) . The principal property of typical elements is that all properties asserted on them are inherited by the members of their corresponding sets. It is important not to confuse the concept of a typical element with the standard concept of “prototype”, which allows for defeasibility, i.e., properties that are not inherited by all of the real members of the set. Asserting a predicate on a typical element of a set is logically equivalent to the multiple assertions of that predicate on all elements of the set. These considerations lead to the distinction between eventuality types and eventuality tokens. The logic defines the following concepts, for which we omit formal details: a. Eventuality types (also known as abstract eventualities): eventualities that involve at least one typical element among their arguments or arguments of their arguments. b. Partially instantiated eventuality types (aka partial instances): a particular kind of eventuality type resulting from substituting the typical elements of some of its (sub-)arguments with other typical elements corresponding to proper subsets. c. Eventuality tokens (also known as instances): a particular kind of partially instantiated even- tuality type with no typical elements in the arguments or sub-arguments11. 10. The formula expresses the de re reading of the sentence, where e1, e2, e3, John, Jack, Ic are first order constants respectively referring to the three eventualities, the two boys, and an ice cream. 11. Actually, ‘instance’ is a term with a broader meaning. There are instances of typical elements that are not eventuali- ties. For simplicity in this paper we assume ‘instances’ and ‘eventuality tokens’ to be synonymous. 23 ROBALDO AND MILTSAKAKI In order to assert that an eventuality e is a, possibly partial, instance of another abstract even- tuality ea, Hobbs introduces the predicate (partialInstance ea e). Another predicate (instance ea e) specifies that e is a total instantiation of ea. It is a consequence of universal instantiation: any property that holds of an eventuality type is true of any (partial) instance of it. We omit here the axioms that formally assert is-a inheritance between eventuality types and their instances. Every relation on eventualities, including logical operators, causal and temporal relations, and even tense and aspect, may be reified into another eventuality. For instance, by asserting (imply′ e e1 e2), we reify the implication from e1 to e2 into an eventuality e and e is, then, thought of as “the state holding between e1 and e2 such that whenever e1 really exists, e2 really exists too”. On the other hand, negation is represented as (not′ e1 e2): e1 is the eventuality of e2’s not existing. The predicates imply′ and not′ are defined to model the concept of ‘inconsistency’. In the next subsection, we show how this concept can be used to construct a uniform account of ‘Direct’ and ‘Indirect’ Contrast. Two eventualities e1 and e2 are said to be inconsistent if and only if they (respectively) imply two other eventualities e3 and e4 such that e3 is the negation of e4. The definition is as follows12: (32) (forall (e1 e2) (iff (inconsistent e1 e2) (and (eventuality e1) (eventuality e2) (exists (e3 e4) (and (imply e1 e3) (imply e2 e4)(not’ e3 e4)))))) The concept of reification used in Hobbs’ logic is suitable for the study of the semantics of discourse connectives because it allows focusing on their meaning while leaving underspecified details about the eventualities involved. In the case of Concession, this amounts to identifying the two eventualities that respectively create and deny the expectation in Argc/Argd , and define the semantics of concessive relations on them. (32) is an example of ‘axiom schema’. In this logic, an ‘axiom schema’ provides one or more different axioms for each predicate p. Axioms determine the expressivity and the computational complexity of the logic. However, the axioms defined in the current version of the logic do not guarantee that the logic is recursively enumerable or computationally tractable. In a real system, we envision handling this problem by defining ontologies for specific domains and making queries to these domains. In what follows, we will briefly illustrate three basic concepts from Hobbs’ logic that we utilize in our proposed semantics of Concession, namely Causality, Defeasible Implication and Likelihood. 7.1.1 CAUSALITY Hobbs’ logic adopts a defeasible account of Causality, originally proposed in Hobbs (1993). This distinguishes between the monotonic notion of ‘causal complex’ and the non-monotonic, defeasible notion of ‘cause’. As Hobbs (1993) explains, when we flip a switch to turn on a light, we say that flipping the switch “caused” the light to turn on. But for this to happen, many other factors need to be satisfied: the bulb is good, the switch is connected to the bulb, there is power in the city, etc. The set of all the states and events that are necessary for the event e to take place as a result, are 12. Hobbs defined several axioms to determine the semantics of the predicates. Those axioms make use of some meta- operators, e.g. if(f), exists, and forall. 24 CORPUS-DRIVEN SEMANTICS OF CONCESSION called the ‘causal complex’ of e. In a causal complex, the majority of participating eventualities are normally true and therefore presumed to hold. In the light bulb case, it is normally true that the bulb is not burnt out, the wiring is in good condition and the power is on, so the conditions are presumed to hold. What cannot be presumed to hold is whether the switch is on or off. Eventualities that cannot be assumed to be true under normal contexts are commonly identified as causes (cf. Kayser and Nouioua (2009)). Based on these ontological grounds, Hobbs represents Causality in terms of two predicates: (cause′ c e1 e2) and (causalComplex s e2). The predicate cause says that c is the state hold- ing between e1 and e2 such that the former is a non-presumable cause of the latter. The predicate causalComplex says that s is the set of all presumable or non-presumable eventualities that are involved in causing e2, including e113. In order to preserve defeasibility, the real existence of the effect e2 does not depend on the real existence of the cause e1 and the causal rule c. In other words, the truth of (Rexist e1) and (Rexist c) does not imply that (Rexist e2) is also true. (Rexist e2) is true just in case all the eventualities in the causal complex of c2 really exist, as asserted by the following axiom: (forall (s e) (if (and (causalComplex s e) (forall (e1)(if (member e1 s) (Rexist e1))) ) (Rexist e) )) It must be pointed out that in practice we can never specify all the eventualities in a causal complex. For instance, consider the following toy example of Concession: (33) Although [John studied hard]Argc , [he did not pass the exam ]Argd . In (33) the expectation “John passed the exam” is created by a defeasible general causal rule “studying hard causes passing exams” that instantiates on the present context. Nevertheless, John did not pass the exam. There was a particular unknown reason why he did not, despite his hard studying. Determining all context-dependent co-causes that had to be in place in order to properly trigger the causal rule would be clearly impossible. Of course, this amounts to saying that NL sentences may be properly interpreted even if causal complexes are unknown. Therefore, to conclude, in most cases the causal complex exists, but it is not possible to infer it. 7.1.2 DEFEASIBLE IMPLICATION Defeasibility does not hold only for causal rules. Most of our everyday knowledge is non-monotonic, i.e., only approximately correct. For example, knowing that birds fly allows us to infer that if Tweety is a bird, then Tweety can fly. This conclusion will be defeated later when we learn that Tweety is actually a penguin and therefore does not fly. The example illustrates that we need to be careful about how we model knowledge of the world. Hobbs, following McCarthy (1980), models com- mon sense implication via monotonic implication (meta-operator if), but allows for defeasibility via the introduction of the underspecified predicate etc in the antecedent of the implication. 13. This is asserted by the following axiom: (forall (e1 e2) (if (cause e1 e2) (exists (s) (and (causalComplex s e2) (member e1 s))))) 25 ROBALDO AND MILTSAKAKI (forall (x) (if (and (bird x) (etc)) (fly x))) The formula says that if x is a bird and has other unspecified properties encoded as etc (i.e., x’s wings are robust enough), then x can fly. In other words, the formula describes the prototype of bird with respect to the property of flying. etc is a conjunction of eventualities that are true for the prototype and allow for the property of flying. For non-flying birds, at least one of those properties does not hold. Although etc is left underspecified in the formulae, its precise definition depends, and so needs to be indexed, on the corresponding predicate, e.g., bird, in the example above. In order to set up a uniform formal account of Concession, we need to introduce a new predicate that denotes non-monotonic implications. Let us term this new predicate as ‘nonMonotonicIf ’. The predicate nonMonotonicIf must have the same syntactic structure as the predicate cause: it must relate two eventualities e1 and e2, and it may be reified into a new eventuality. As for e1 and e2 they can be abstract eventualities or instances. Obviously, (nonMonotonicIf e1 e2) is true iff e1 defeasibly implies e2. The definition of nonMonotonicIf is reported in (34). An eventuality e1 defeasibly implies e2 if and only if for each partial instance of e1 there is a partial instance of e2 for which the meta-predicate if, augmented with an opportune etc predication, holds. Of course, if and only if e1 and e2 are two instances, it is necessary to assume that the predicate partialInstance denotes a reflexive relation, i.e., that every eventuality is a partial instance of itself. p1 and p2 are the predicates indicating the types of the eventualities e1 and e2 respectively. Hobbs defines a meta-predicate pred to relate an eventuality with the unique predication that describes it: (pred p e) states that p is the predicate whose reification is e. (34) (forall (e1 e2) (iff (nonMonotonicIf e1 e2) (forall (e ′1) (if (partialInstance e ′1 e1) (exists (e ′2) (and (partialInstance e ′2 e2) (forall (x1 x2 . . . xn) (if (and (pred p1 e ′1)(p1 x1 x2 . . . xn)(etc)) (and (pred p2 e ′2)(p2 x1 x2 . . . xn))))))))) ) 7.1.3 LIKELIHOOD Eventualities exist in a Platonic universe of possible individuals: entities, states and events. As said above, if they happen to actually occur in the real world, that is one of their properties, and we express it with the predicate Rexist. Real existence is one mode of existence but there are others, too. The eventuality could be part of someone’s beliefs but not occur in the real world. It could be merely possible or likely but not real. It could, also, be unlikely or impossible. An especially important modality is “happening at a particular time”. Possibility is one common judgment we make about eventualities in situations of uncertainty. Likelihood is another. Likelihood is intended as the common sense notion of the mathematical version of probability. Mathematically defined probability is a special case of common sense likelihood. Likelihood is a qualitative notion intended to model the vague probability judgements we make in everyday life, as when we say that it’s likely to rain or that the train may be late. 26 CORPUS-DRIVEN SEMANTICS OF CONCESSION Likelihoods are members of a partially ordered scale of likelihoods. For Hobbs, such a scale s satisfies the predicate (likelihoodScale s). The likelihood of an eventuality e is with respect to an implicit set of constraints c defining the sample space. An eventuality c may be defined as a single eventuality ec that reifies the conjunction of all the constraints such that (and ′ ec e1 . . . en) is true, where e1, . . . , en are the eventuality-constraints. The likelihood of e is given in the context where the predicate (Rexist ec) is assumed to hold. With the formula (likelihood d e c), where d is a number, e an eventuality, and c a set of con- straints, we assert that d is the likelihood of e’s really existing, if the set of eventualities in c really exist and d belongs to the contextually relevant likelihood scale s. We say that a certain eventuality e is ‘likely’ when a set of eventualities c holds, iff the likelihood of e given c is a qualitative value belonging to the highest part of the contextually relevant likelihood scale. (forall (e c) (iff (likely e c) (exists (s d s1) (likelihood d e c) (likelihoodScale s) (belong d s1) (high s1 s)) )) Likelihood is connected to other modalities via additional axioms. If the likelihood of an even- tuality e with constraints c is the top of the likelihood scale, then e is necessary given c, i.e., it is implied from the latter. If the likelihood of e is the bottom of the likelihood scale, then it is not possible given c. In the next subsection, we use the above definitions of Causality, non-monotonic Implication and Likelihood to define the semantics of Concession with respect to the different sources of expectation that we identified in our analysis in the PDTB corpus. 7.2 A refined logical account of Concession In this section, we propose logical formulae in Hobbs’ logic that represent the meaning of concessive relations using a uniform representation that is minimally adjusted to reflect the interpretation of the different sources of expectation. As discussed in Section 1, previous approaches of Concession, e.g., Winter and Rimon (1994) and Lagerwerf (1998), mostly focus on how the expectation is denied. The distinction between ‘Denial of expectation’ and ‘Concessive opposition’ is in this spirit. In ‘Denial of expectation’, Argd directly denies the expectation. In ‘Concessive opposition’, Argd entails the negation of the expectation. In this line of work, not a lot of attention has been paid to how Argc creates the expectation. In most cases, it is simply assumed that the expectation is created from Argc via an underspecified entailment ‘→’. Our approach is different in that it focuses on characterizing how the expectation is created rather than how it is denied. In formal terms, we propose to represent semantic concessive relations via the following general pattern in Hobbs’ logic: (35) (exist (sac sc ec ee ed) ∧ (partialInstance sc sac) ∧ (Φ sc ec ee) ∧ (Rexist ec) ∧ (Rexist ed) ∧ (inconsistent ee ed) ) 27 ROBALDO AND MILTSAKAKI where Φ is a generic predicate referring to the underspecified entailment that creates the expectation; it corresponds to Winter&Rimon’s ‘→’. According to our analysis, Φ can be any of the predicates cause’, nonMonotonicIf’, or likely’ as defined by Hobbs’ account and illustrated in the previous section. The eventuality ee corresponds to the created expectation. It is created from the eventuality ec, conveyed by Argc, via the relation Φ. The eventuality ed is conveyed by Argd. The formula in (35) asserts that ee and ed are inconsistent, via the predicate (inconsistent ee ed). According to the definition in (32), whether the inconsistency is achieved directly or indirectly remains underspecified. (inconsistent ee ed) comes out true iff ee implies a third, possibly different, eventuality, and ed its negation. Obviously, it may be either the case that ed is already the negation of ee (Denial of expectation), or that it implies it (Concessive opposition). The eventuality sac is a general abstract (defeasible) rule. The formula asserts that the reification of Φ, i.e. the eventuality sc, is a more specific instantiation of sac . Both ec and ed really exist in the context, as asserted in (35) via the predicate Rexist . On the contrary, sac and sc do not necessarily exist in the real world; as exemplified below in (45), they could exist only in the speaker’s beliefs. The general pattern in (35) covers the cases of Causality, non-monotonic Implication, and Correlation. In (36) we repeat the toy examples used earlier to demonstrate the three semantic classes. (36) a. Although [John studied hard]Argc , [he did not pass the exam ]Argd . (Causality) b. [Penguins are birds ]Argc . Nevertheless [they do not fly]Argd . (Implication) c. [John will do his report ]Argd , but [he will do it at home]Argc . (Correlation) The formulae in Hobbs’ logic that represent examples (36.a-c) are shown in (37), (38), and (39) respectively. Note that, with respect to the general pattern in (35), the three formulae below differ only in the predicate Φ. In (37), it has been substituted by cause’, in (38) by nonMonotonicIf’, and in (39) by likely’. In the next subsection, we will show that (35) is also able to account for the generalizations identified above in Section 4. (37) (exist (sac sc ec ee ed) (partialInstance sc sac) ∧ (cause’ sc ec ee) ∧ (Rexist ec) ∧ (Rexist ed) ∧ (inconsistent ee ed) ) ec = “John studied hard” sac = “Studying hard causes passing exams” ee = “John passed the exam” ed = “John did not pass the exam” 28 CORPUS-DRIVEN SEMANTICS OF CONCESSION (38) (exist (sac sc ec ee ed) (partialInstance sc sac) ∧ (nonMonotonicIf’ sc ec ee) ∧ (Rexist ec) ∧ (Rexist ed) ∧ (inconsistent ee ed) ) ec = “Penguins are birds” sac = “Birds fly” ee = “Penguins fly” ed = “Penguins do not fly” (39) (exist (sac sc ec ee ed) (partialInstance sc sac) ∧ (likely’ sc ec ee) ∧ (Rexist ec) ∧ (Rexist ed) ∧ (inconsistent ee ed) ) ec = “John will do his report” sac = “John does not usually do his reports at home” ee = “John will not do his report at home” ed = “John will do his report at home” 7.3 Abductive usage of defeasible rules Lagerwerf (1998) identified three different sub-cases of Concession. The three examples given in (36) are cases of ‘Denial of expectation - Content’, according to Lagerwerf’s terminology. The other two subcases are ‘Denial of expectation - Epistemic’ and ‘Denial of expectation - Speech Act’. For convenience, we give the examples again in (40): (40) a. [Mary loves you very much ]Argd , although [you already know that]Argc . (Denial of expectation - Speech Act) b. [Theo was not exhausted ]Argd , although [he was gasping for air]Argc . (Denial of expectation - Epistemic) In (40.a) the expectation is denied by the illocution of Argd, i.e., its Speech Act, rather than by its locutionary meaning. It is the fact that I tell you Argd, and not Argd, that is inconsistent with the expectation, created by the rule “If I know that you already know something, I won’t’ say it to you again”. The formalization of occurrences of Concession involving Speech Acts are simple in Winter and Rimon (1994) and Lagerwerf (1998). The default implication is asserted on the Speech Act associated with the proposition rather than the proposition itself. Speech Acts do not raise any problems in our approach, either, and so we will not go into the details. Consistent with Hobbs’s logic, the Speech Act of an eventuality is simply reified into a new eventuality, and the defeasible rules are asserted on the latter. The interesting case is ‘Denial of expectation - Epistemic’, shown in (40.b). In this example, there is a causal rule “Being exhausted causes gasping for air”, and the expectation is created abduc- 29 ROBALDO AND MILTSAKAKI tively, i.e., by observing that Theo was gasping for air, it may be concluded that he was exhausted. Lagerwerf (1998) formalizes this intuition in Predicate Logic as follows14: ∀x[Gfb(x) > B(i, Exh(x))] where Gfb and Exh are predicates denoting the set of individuals gasping for air and the set of exhausted individuals, respectively. B(y,Φ) is an epistemic operator asserting that y believes Φ, i refers to the speaker, and ‘>’ is the defeasible implication operator defined in Asher and Morreau (1991). As discussed in Section 4, Lagerwerf’s intuition is correct in our view, but the proposed formalization seems to deviate from the heart of the intuition. As discussed above in (18.a), the defeasible rule is general and therefore does not apply specif- ically to any particular speaker. Therefore, in the formalization, i should be most properly sub- stituted by a universal quantification over all possible believers. Similar remarks may be found in Pander Maat (1998), who argue that three different perspectives of subjectivity may be ascribed to the belief of the statements in a discourse: ‘objective’, ‘speaker’, and ‘other’ perspective. The latter holds when the belief of a statement is ascribed to people other than the speaker. Consider the following examples of ‘Denial of expectation’ taken from the PDTB15: (41) a. Although [it adopted a poison-pill defense]Argc , [the board believed that Mr. Icahn is more interested in talking the stock price higher than acquiring USX]Argd . (Causality; Denial of expectation - Epistemic) b. [Sure, price action is volatile and that’s scary ]Argc . But [all-in-all stocks are still a good place to be ]Argd . (Implication; Denial of expectation - Epistemic) c. [Program trading increases volatility ]Argc , but [I don’t think it should be banned ]Argd . (Causality; Denial of expectation - Content) In (41.a), by observing that it adopted a particular strategy, the board, and not the speaker, should believe that Mr. Icahn is interested in acquiring USX. On the other hand, in (41.b), just like in (40.b), the statements and the conclusions that may be drawn are presented as true objective facts: by observing that price action is volatile everyone should conclude that all-in-all stocks are no longer a good place to be. Finally, (41.c) highlights that the belief of a statement may be ascribed to the speaker also when the defeasible rule is used deductively (Lagerwerf’s ‘Denial of expectation - Content’). In (41.c), by observing that program trading increases volatility, the speaker, but not necessarily everybody else, believes that it must be banned. It should be made clear, however, that the assertion of the proper perspectives is orthogonal to the semantics of Concession. Even the general defeasible rules may, in principle, be asserted in the speaker’s beliefs only. Consider the following toy example: (42) Although [I watered my feet every morning for one month]Argc , [I did not get taller ]Argd . (Causality) 14. On the other hand, as shown above in (14.b), the defeasible rule used in (40.a) could be formalized as: ∀x[K(i, K(y, x)) >� ¬T(i, y, x)], where K(a, x) is an epistemic operator asserting that a knows x, T(a1, a2, x) is true if a1 tells x to a2, i and y are two constants respectively referring to the speaker and the hearer. 15. We did not correct typos that appeared in the corpus. (41.a) should be corrected to “. . . Mr. Icahn is more interested in taking the stock . . .” . 30 CORPUS-DRIVEN SEMANTICS OF CONCESSION From the example in (42), we understand that the speaker believes that “watering the feet causes growing up”. In Hobbs’, such an abstract causal rule is reified into an eventuality sac. Then, in order to ascribe it to the speaker’s subjectivity only, a separate predicate like (B i sac) can be conjoined to the whole formula. In other words, in all other examples seen above, each defeasible rule has been always taken as a true fact, but obviously in the case that the hearer does not believe it, it may be asserted as a speaker’s belief only. In our view, Lagerwerf (1998)’s intuition must be formalized exactly as it is stated. In (40.b), the defeasible causal rule is “being exhausted causes gasping for air”, and it yields the expectation abductively. Thus, the formula is the one in (43); the only difference with respect to the formulae associated above with occurrences of Concession where the expectation is created via Causality is the assertion of (cause’ sc ee ec) in place of (cause’ sc ec ee). (43) (exist (sac sc ec ee ed) (partialInstance sc sac) ∧ (cause’ sc ee ec) ∧ (Rexist ec) ∧ (Rexist ed) ∧ (inconsistent ee ed) ) ec = “Theo was gasping for air” sac = “Being exhausted causes gasping for air” ee = “Theo was exhausted” ed = “Theo was not exhausted” If the causes/effects of the causal rule or the causal rule itself are believed by the speaker or by someone else, this is separately asserted via additional conjuncts. The formulae corresponding to (41.a) and (42) are shown in (44) and (45) respectively (in the formulae, b is a constant referring to the board and i is a constant referring to the speaker). (44) (exist (sac sc ec ee ed) (partialInstance sc sac) ∧ (cause’ sc ee ec) ∧ (Rexist ec) ∧ (B b sac) ∧ (B b ed) ∧ (inconsistent ee ed) ) ec = “It adopted a poison-pill defence” sac = “Mr. Icahn’s intention of acquiring USX causes adopting a poison-pill defence” ee = “Mr. Icahn wants to acquire USX” ed = “Mr. Icahn does not want to acquire USX” (45) (exist (sac sc ec ee ed) (partialInstance sc sac) ∧ (cause’ sc ec ee) ∧ (Rexist ec) ∧ (Rexist ed) ∧ (B i sac) ∧ (inconsistent ee ed) ) ec = “I watered my feet” sac = “Watering the feet causes getting taller” ee = “I got taller” ed = “I did not get taller” 31 ROBALDO AND MILTSAKAKI 8. Conclusions 8.1 Sources of expectation We presented an empirical analysis of Concession based on the annotations of Concessive connec- tives in the Penn Discourse Treebank. In concessive relations, one argument gives rise to an expectation which is then denied in the second argument. In this paper, we have argued that a proper account of Concession should be grounded on how the expectation is created. Specifically, we identified four sources of expecta- tion: Causality, Implication, Correlation, and Implicature. In Causality, the created expectation is causally related to the eventuality that creates it. In Implication, the expectation is created on the basis of a specific property associated with the eventuality that creates the expectation and in Cor- relation the expectation is related to the eventuality that creates it via co-occurrence. We termed “Implicature” the source of expectation that involves pragmatic inferencing. We leave a formal account of this semantically complex category for future work. To test the reliability of these categories, we conducted an inter-annotator agreement study on one thousand Concessive tokens in the PDTB. The high kappa score confirms that the categories can be distinguished reliably. To evaluate the practical merits of our approach over previous accounts of Concession, we con- ducted a second inter-annotator study. In this study, two linguistic students annotated 200 instances from PDTB that had been annotated as Concession by one annotator and Contrast by the other. We trained the new annotators with the new definition of Concession, explaining the four sources of expectations. There was more than 80% agreement in this dataset which is a very significant improvement in making reliable distinctions between Contrast and Concession. 8.2 Formal treatment of Concession Following earlier work by Lagerwerf (1998), we refined the semantics of Concession using basic constructs from the logic proposed in Hobbs (1998). Central to Hobbs’ proposal is the notion of reification which allows complex natural language statements to be modelled in Predicate Logic. We propose that every type of Concession presupposes a general rule that holds in the context. We propose a logical account that models the semantics of Concession by defining the rule for each type of Concession. In the case of Causality, we infer a defeasible causal relationship between the eventuality expressed in one argument and the eventuality of the expectation. In Implication, we infer that the eventuality expressed in one argument non-monotonically entails the eventuality of the expectation. In Correlation we infer that the eventuality expressed in the argument that creates the expectation is likely to co-occur with the eventuality of the expectation. The proposed logic formulae differ by the kind of predicate describing the abstract rule (cause’, nonMonotonicIf’, likely’). 8.3 Impact and future work We have shown that by identifying the different sources of expectation in Concession not only are we able to characterize more precisely the events that give rise to expectations but we obtain more reliable semantic distinctions between Concession and Contrast. We were able to obtain empirical evidence for this claim by a new inter-annotator study that showed significant inter-annotator agree- 32 CORPUS-DRIVEN SEMANTICS OF CONCESSION ment for tokens that were previously confusing (tokens of annotator disagreement between Contrast and Concession). The proposed account of Concession and its logical treatment is an improvement over some previous accounts which are insufficient in capturing the range of variants identified in naturally occurring data. We maintain that the identified sources of expectation in Concession and their logi- cal treatment adequately demonstrated how the study of discourse relations as attested in naturally occurring data can help us improve our understanding of the semantics of discourse relations. An important aspect of concession, the source of expectation, was overlooked in previous approaches but became apparent when studying the data. Thus, addressing our second question “What kind of semantic representation will allow covering the rich range of variants conveying Concession and Contrast?”, we concluded that defining pred- icates which take as arguments reified eventualities (and even speech acts) is critical for handling discourse level semantic relations. Hobbs’s proposed semantics for natural language has proven to be especially well suited for articulating a uniform model of Concession while accounting for the range of variants in a simple and straightforward manner. For future work, we need to delve deeper in the cases in which the created expectation involves pragmatic reasoning. Moreover, we need to further test the ambiguity between Contrast and Con- cession that, in the present paper, was studied only with respect to 200 tokens that were identified as problematic in the PDTB. Finally, although the identification of the four sources of Concession was empirically tested against a larger set of occurrences (1000 PDTB tokens), in order to see how our results are generalizable, we advocate further experiments on corpora pertaining to different genres and in other languages. While we believe that our approach is in the right direction, defining an important step to- wards processing discourse problems automatically, the proposed semantics cannot be readily im- plemented in current state-of-the-art inference systems. Significantly more work would be required to integrate the proposed semantics in a real system, e.g., TACITUS system Hobbs (1986), Montaz- eri and Hobbs (2011), Ovchinnikova et al. (2011), which implements Hobbs’ logic. On the other hand, most current systems are based on shallow features. A recent proposal along this line is the one of Meyer and Popescu-Belis (2012). They trained a statistical classifier on PDTB data for disambiguating discourse connectives, among which discourse connectives convey- ing Concession and Contrast. The classifier involves a large set of syntactic and semantic features and it is used for enhancing the performances of a separate Statistical Machine Translation system. Classifiers based on shallow features could also benefit from our work. For instance, the overall per- formances of Meyer and Popescu-Belis (2012)’s classifier could perhaps be improved by including semantic features specifically aimed at identifying the sources of concessive relations. References W. Abraham. Discourse particles in german: how does their illocutive force come about? In W. Abraham, editor, Discourse particles in German, pages 203–252. John Benjamins, Amster- dam/Philadelphia, 1991. J.C. Anscombre and O. Ducrot. Deux mais en francais? Lingua, 43:23–40, 1979. N. Asher. Reference to Abstract Objects in Discourse. Dordrecht, Kluwer, 1993. 33 ROBALDO AND MILTSAKAKI N. Asher and M. Morreau. Commonsense entailment: a modal theory of nonmonotonic reasoning. In Proc. of the 12th international joint conference on Artificial intelligence - Volume 1, pages 387–392, San Francisco, CA, USA, 1991. Morgan Kaufmann Publishers Inc. E. Bach. On time, tense, and aspect: An essay in english metaphysics. In P. Cole, editor, Radical Pragmatics, pages 63–81. Academic Press, New York, 1981. Diane Blakemore. Denial and contrast: A relevance theoretic analysis of ‘but’,. Linguistics and Philosophy, 12(1):15–37, 1989. J. Carletta. Assessing agreement on classification tasks: the kappa statistic. Computational Lin- guistics, 22:249–254, 1996. Lynn Carlson, Daniel Marcu, and Mary Ellen Okurowski. Building a discourse-tagged corpus in the framework of rhetorical structure theory. In Proceedings of the Second SIGdial Workshop on Discourse and Dialogue - Volume 16, SIGDIAL ’01, pages 1–10, Stroudsburg, PA, USA, 2001. Association for Computational Linguistics. M. Dascal and T. Katriel. Between semantics and pragmatics: the two types of but - hebrew ’aval’ and ’ela’. Theoretical Linguistics, 4:143–172, 1977. D. Davidson. The logical form of action sentences. In Nicholas Rescher, editor, The Logic of Decision and Action. University of Pittsburgh Press, 1967. B. Di Eugenio. On the usage of kappa to evaluate agreement on coding tasks. In In Proceedings of the Second International Conference on Language Resources and Evaluation, pages 441–444, 2000. A. Foolen. Polyfunctionality and the semantics of adversative conjunctions. Multilingua, 10-12: 79–92, 1991. N. Francez. Contrastive logic. Logic Journal of the IGPL, 3(5):725–744, 1995. P. Grice. The causal theory of perception. In Proc. of the Aristotelean Society, Vol. 35, pages 121–152, 1961. P. Grice. Logic and conversation. In P. Cole and J. L. Morgan, editors, Syntax and Semantics: Vol. 3: Speech Acts, pages 41–58. Academic Press, San Diego, CA, 1975. B. Grote, N. Lenke, and M. Stede. Ma(r)king concessions in english and german. Discourse Processes, 24(1):87–117, 1995. J. R. Hobbs. On the coherence and structure of discourse. Technical Report 85-37, Stanford, CA: Stanford University, Center for the study of Language and Information, 1991. J. R. Hobbs. Toward a useful notion of causality for lexical semantics. Journal of Semantics, 22: 181–209, 1993. J.R. Hobbs. Overview of the TACITUS project. Computational Linguistics, 12(3), 1986. 34 CORPUS-DRIVEN SEMANTICS OF CONCESSION J.R. Hobbs. Monotone decreasing quantifiers in a scope-free logical form. In in Semantic Ambiguity and Underspecification, pages 55–76. 1995. J.R. Hobbs. The logical notation: Ontological promiscuity. In Chapter 2 of Discourse and Inference. 1998. Available at http://www.isi.edu/∼hobbs/disinf-tc.html. L.R. Horn. A Natural History of negation. Cambridge University Press, Cambridge, 1989. M. Izutsu. Contrast, concessive, and corrective: Towards a comprehensive study of opposition relations. The Journal of Pragmatics, 40:646–675, 2008. D. Kayser and F. Nouioua. From the textual description of an accident to its causes. Artificial Intelligence, 173:1154–1193, 2009. A. Kehler. Interpreting cohesive forms in the context of discourse inference. In Proc. of the 32th Meeting of the Association of Computational Linguistics, 1994. A. Kehler. Coherence, Reference, and the Theory of Grammar. CSLI Publications, 2002. E. König. Concessive, connectives, and concessive sentences: cross-linguistic regularities and prag- matic principles. In J.A. Hawkins, editor, Explaining language universal, pages 145–166. Black- well, London, 1983. I. Korbayova and B. Webber. Interpreting concession statements in light of information structure. In Bunt H. and Muskens R., editors, Interpreting concession statements in light of information structure, pages 145–172. Kluwer Academic Publishers, 2007. L. Lagerwerf. Causal connectives have Presuppositions: Effects on Coherence and Discourse structure. PhD thesis, The Hague: Holland Academic Graphics, The Netherlands, 1998. R. Lakoff. If’s, and’s, and but’s about conjunction. In In C.J. Fillmore and D.T. Langendoen, editors, Studies in Linguistic Semantics. Holt, Reinhart and Winston, New York, 1971. E. Lang. The semantics of coordination. John Benjamins B.V., Amsterdam, 1984. E. Lang. Adversative connectors on distinct levels of discourse: A re-examination of eve sweetser’s three-level approach. In Bernd / Couper-Kuhlen Kortmann, editor, Cause, condition, concession, contrast, pages 235–256. Berlin: Mouton de Gruyter, 2000. W. C. Mann and S. A. Thompson. Rhetorical structure theory: Toward a functional theory of text organization. Text, 8(3):243–281, 1988. J. McCarthy. Circumscription: A form of nonmonotonic reasoning. Artificial Intelligence, 13: 27–39, 1980. J. McCarthy. Epistemological problems of artificial intelligence. In Proc. of the International Joint Conference on Artificial Intelligence, pages 1038–1044, Cambridge, Massachusetts, 2002. T. Meyer and A. Popescu-Belis. Using sense-labeled discourse connectives for statistical machine translation. In Proc. of the Joint Workshop on Exploiting Synergies between Information Retrieval and Machine Translation (ESIRMT) and Hybrid Approaches to Machine Translation (HyTra), pages 129–138, 2012. 35 ROBALDO AND MILTSAKAKI E. Miltsakaki, L. Robaldo, A. Lee, and A. Joshi. Sense annotation in the penn discourse treebank. In Proc. of Computational Linguistics and Intelligent Text Processing, pages 275–286, Cambridge, Massachusetts, 2008. N. Montazeri and J.R. Hobbs. Elaborating a knowledge base for deep lexical semantics. In J. Bos and S. Pulman, editors, In Proc. of 9th International Workshop on Computational Semantics, pages 195–204, 2011. J. Moore and M. Pollack. A problem for RST: The need for multi-level discourse analysis. Compu- tational Linguistics, 18:537–544, 1992. E. Ovchinnikova, N. Montazeri, T. Alexandrov, J.R. Hobbs, M.C. McCord, and R. Mulkar-Mehta. Abductive reasoning with a large knowledge base for discourse processing. In Proc. of the Ninth International Conference on Computational Semantics (IWCS 2011), pages 225–234, 2011. H. Pander Maat. Classifying negative coherence relations on the basis of linguistic evidence. Jour- nal of Pragmatics, 30(2):177–204, 1998. M. Poesio and R. Artstein. Anaphoric annotation in the ARRAU corpus. In Proc. of Language Resources and Evaluation (LREC08), 2008. R. Prasad, E. Miltsakaki, N. Dinesh, A. Lee, A. Joshi, B. Webber, and L. Robaldo. The penn discourse treebank 2.0. annotation manual. Technical Report IRCS-06-01, Institute of Research in Cognitive Science, University of Pennsylvania, 2008. J. Sanders. Perspective in narrative discourse. PhD thesis, Tilburg University, The Netherlands, 1994. T.J.M. Sanders, W.P.M. Spooren, and L.G.M. Noordman. Toward a taxonomy of coherence rela- tions. Discourse Processes, 15(1), 1992. T.J.M. Sanders, W.P.M. Spooren, and L.G.M. Noordman. Coherence relations in a cognitive theory of discourse representation. Cognitive Linguistics, 4(2), 1993. J. Spenader and A. Lobanova. Reliable discourse markers for contrast relations. In Proc. of the Eighth International Conference on Computational Semantics (IWCS-8), pages 210–221, 2009. W. Spooren. Some aspects of the Form and Interpretation of Global Contrastive Coherence. PhD thesis, Nijmegen University, The Netherlands, 1989. E. Sweetser. From Etymology to Pragmatics. Metaphorical and Cultural Aspects of Semantic Struc- ture. Cambridge University Press, Cambridge, 1990. Y. Versley and A. Gastel. Linguistic tests for discourse relations in the tba-d/z treebank of german. Dialogue & Discourse, 4(2):142173, 2013. A. von Klopp. But and negation. Nordic Journal of Linguistics, 17(1), 1994. Y. Winter and M. Rimon. Contrast and implication in natural language. Journal of Semantics, 11: 365–406, 1994. 36