incremental processing of telicity in italian children jasmijn e. bosch, mathilde chailleux & francesca foppolo* abstract. a sentence like ‘lyn has peeled the apple’ triggers a telicity inference that the event is telic and to a culmination inference that the event has reached its telos and has stopped. this results in the final interpretation of the sentence that lyn has completely peeled the apple. we present an eye-tracking study to test children’s ability to predict the upcoming noun (e.g., the apple) during the incremental processing of sentences like ‘show me in which picture lyn has peeled the…’ in which the predicate is telic and the verb appears in the perfective form. by means of the visual world paradigm, our aim was to investigate children’s ability to use the lexical semantics and aspectual morphology of verbs during language processing and comprehension. to test if children can predict the target (e.g., a completely peeled apple) by exploiting the lexical-semantic meaning of the verb, we contrasted the picture of the target with the picture of an object that cannot be peeled; to test if they can predict on the basis of the verb’s perfective morphology, we compared the target with the picture of a half-peeled apple. our results show that italian children anticipate the upcoming noun in both cases, providing evidence that they can incrementally exploit the morphosyntactic cue on the verb (perfective aspect) to derive the culmination inference that the telos is reached, and the event is completed. we also show that the integration of aspect requires some additional time compared to the integration of basic lexical semantics of the verb. keywords. aspect, semantics, anticipation, eye-tracking, visual world paradigm 1. introduction. it is a well-known fact that, during sentence processing, we can anticipate upcoming nouns on the basis of the lexical semantics of prenominal verbs. in their pioneer study, altmann & kamide (1999) showed that participants anticipated the word ‘cake’ when hearing the verb ‘eat’, which was revealed by increased looks towards the picture of a cake in a visual scenario in which the cake was the only edible object. other studies have shown that we can also integrate morphosyntactic information, such as tense expressed on the verb, to anticipate the target referent. in another study employing the visual world paradigm, altmann & kamide (2007) showed that, when participants heard a sentence like ‘the man will drink…’, they looked more at a full glass of beer (in which its content is yet to be drunk), while they looked more at the empty glass of wine when hearing the sentence ‘the man has drunk…’. this suggests that, as listeners, we quickly and incrementally integrate the inference conveyed by the verb’s morphology that an event of drinking has taken place and was completed, resulting in an empty glass. in fact, there is more than lexical information and tense in a verb. verbs denote events, and events are complex phenomena. they stretch along the time dimension and can be classified as punctive or durative. for example, events such as ‘blowing out a candle’, which happen immediately, are called punctive and events like ‘peeling an apple’, which take some time, are called durative. independently of their duration, predicates can also have a telos (or a culmination point); this property, however, is not (or not always) intrinsic to the lexical properties of the predicate itself, *this project has received funding from the european union’s horizon 2020 research and innovation programme under the marie skłodowska curie grant agreement no 765556. the affiliation of all authors is the university of milano-bicocca. contact emails: jasmijn.bosch@unimib.it, mathilde.chailleux@gmail.com, francesca.foppolo@unimib.it. proceedings of elm 1: 071-077, 2021 c©2021 jasmijn e. bosch, mathilde chailleux and francesca foppolo published by the lsa with permission of the author(s) under a cc by license. 71 https://doi.org/10.3765/elm https://www.elm-conference.net/ but it pertains to the combination of a certain predicate (e.g., ‘peeling’) with a certain complement. this is called the aktionsart. considering this property of events, predicates can be classified as atelic or telic. under this distinction, predicates like, for example, ‘running’ or ‘peeling apples’ are classified as atelic, since they do not have a clear ending or culmination point. on the other hand, predicates like ‘running the boston marathon’ or ‘peeling an apple’ are classified as telic, due to the fact that there is a clear point in which the event culminates (namely, when the finish line in the marathon is reached or when the apple is completely peeled). there is another layer to add to the time dimension of the processing of events: in many languages, verbs carry tense and aspectual morphology. while tense information deals with the collocation of the event in time along a linear dimension of past-present-future, aspect deals with the status of the completion of the event. with respect to this distinction, a verb like ‘is peeling’ carries the imperfective aspectual information, signaling the fact that the event of peeling has started at some point in the past, but is still ongoing. in the case of telic predicates like ‘is peeling the apple’, this means that the action of peeling is still happening, and the telos has not been reached yet. when combined with perfective aspect, instead (i.e., ‘has peeled the apple’), the information conveyed is that the event of peeling has stopped, and the telos has been reached. thus, in a simple sentence like ‘lyn has peeled the apple’, two different layers of information are carried by the verb and are integrated online during sentence comprehension. the first one is the layer of the aktionsart, which carries the information that the event is durative and telic (since the durative predicate is combined with a definite noun phrase, ‘the apple’). the second layer consists of the aspectual morphology realized on the verb, which carries the information about the status of completion of the event. together, these two layers of information point to a telicity inference that the event is telic and to a culmination inference that the event has reached its telos and has stopped. this results in the final interpretation of the sentence that lyn has completely peeled the apple. the nature and interplay of these inferences in the derivation of the meaning of the sentence are debated in theoretical semantics, and they go beyond the purposes of this paper. from a psycholinguistic perspective, the question that can be tested experimentally is whether (and when) this interpretation is carried out incrementally during sentence processing. previous works have focused on the time course of the telicity inference. for example, proctor et al. (2004) used a self-paced reading study to examine the speed and accuracy with which readers draw telicity inferences during on-line language comprehension. participants read sentences containing either a consumption verb (e.g., ‘consume’) or an observation verb (e.g., ‘monitor’) followed by either a mass or a count object (e.g., ‘ice water’ vs. ‘ice cube’), which could trigger an atelic or telic interpretation of the predicate. each sentence ended with an adverbial phrase that was either consistent (e.g., ‘in 8 minutes’) or inconsistent (e.g., ‘for 8 minutes’) with a telic verb. proctor and collaborators report no slowdown in the final adverbial region in sentences like ‘leslie consumed polar purity’s ice cube for eight minutes’, while a slowdown was observed in the same sentence when the adverbial modifier ‘in eight minutes’ was used. they interpret this result as evidence that the inference that an event is telic seems not to be computed until there is evidence that it is needed and that, when it is computed, this inference is associated with a computational cost (see also, pickering et al. 2006, townsend 2012). few works have used the visual world paradigm to test the incremental processing of aspect. among these, zhou, crain & zhan (2014) tested mandarin-speaking adults and children aged 3 to 5 with sentences in which the verb carried a perfective aspectual morpheme (le) or a durative aspectual morpheme (zhe) in a scenario contrasting a completed vs. an uncompleted event. the results show that both the adults and the children (of all age groups) looked more at the completed proceedings of elm 1: 071-077, 2021 jasmijn e. bosch, mathilde chailleux and francesca foppolo: incremental processing of telicity in italian children. 72 https://doi.org/10.3765/elm https://www.elm-conference.net/ event (e.g., a woman who has finished planting a flower) when hearing the perfective morpheme, while they looked more to the ongoing event (e.g., a woman in the process of planting a flower) when hearing the durative morpheme. this effect occurred immediately after the onset of the aspectual morpheme (appearing at the end of the verb), showing that even young children are able to use the temporal information encoded in aspectual morphemes as rapidly as adults to facilitate event recognition. summing up, we know that we process lexical and tense features of verbs incrementally, and that we make use of such cues to identify objects that are compatible with the lexical semantics of the verb and to separate completed from yet-to-be events (as shown in altmann and kamide’s studies mentioned above). we also know that we make use of aspectual cues to separate ongoing vs. completed events, as shown by zhou and colleagues. previous reading studies, instead, suggest that we do not immediately commit to the fact that the event is telic and has a culmination point, even when we encounter a durative predicate followed by a definite noun phrase. however, some questions still remain unanswered. suppose that we already know, from the visual context, that we are facing a telic event such as peeling an apple, and we are shown different degrees of completion of the same peeling-an-apple event (e.g., a half peeled apple and a completely peeled apple). in such a case, it is unclear when listeners commit to the fact that the telos is reached, and consequently, when they start looking at the completely peeled apple while hearing a sentence like ‘show me where lyn has peeled the…’. this question has been addressed experimentally by foppolo, greco, panzeri and carminati (2016), who tested italian adults with a visual world paradigm. results show that participants fixated on the completed event while hearing the verb, providing evidence for an incremental and rapid integration of aspectual cues during sentence processing. building on this study, we tested incremental processing of verbs in italian-speaking children, by contrasting the integration of lexical-semantic information and aspectual cues during sentence processing. 2. methods. 2.1 participants. we tested 35 monolingual italian children between the ages of 8 and 10. the data of six children had to be excluded because of poor calibration of the eye-tracker. the final sample consisted of 29 children (11 boys and 18 girls), with a mean age of 9 years and 4 months (sd = 11 months). participants were recruited at a primary school in the urban area of milan. prior to testing, all parents signed a consent form that was approved by the ethics committee of the university of milano-bicocca. 2.2 materials. we created a visual world eye-tracking experiment, in which we tested to what extent children were able to rely on aspectual and lexical information during online sentence processing. the experiment was implemented in e-prime 3 (psychology software tools, pittsburgh, pa). participants saw a visual scenario with two pictures depicting completed or ongoing events. the pictures were colored photographs focusing on the hands of a person who was involved in an action or who had just completed an action using one or more objects. no faces were shown. we used partly the same pictures as foppolo et al. (2016), although additional stimuli were created for the purpose of this study. the two images (577 x 408) were shown on the left and right side of a grey screen (1920 x 1080), with a clear space in between them. regarding the auditory stimuli, participants listened to transitive sentences in italian with a verb in the passato prossimo, in which the auxiliary ha was combined with a past participle to trigger the completion interpretation. for example, participants heard guarda in quale foto ha colorato la stella (‘look in which picture (he/she) colored the star’). three different experimental conditions were proceedings of elm 1: 071-077, 2021 jasmijn e. bosch, mathilde chailleux and francesca foppolo: incremental processing of telicity in italian children. 73 https://doi.org/10.3765/elm https://www.elm-conference.net/ created: two in which the target could be anticipated while processing the verb and one in which no anticipation was possible (these are labelled early and late respectively). in the early lexical condition, the correct picture could be selected on the basis of the lexical meaning of the verb (e.g., an object that can be colored, like a drawing, vs. an object that cannot, like a lego tower, cf. figure 1, a-b). in the early aspect condition, the pictures displayed two actions at a different state of completion (e.g., a half-colored star vs. a fully colored star, cf. figure 1, a-c), so that the correct picture could be selected based on the aspectual information morphologically expressed by the verb that should trigger the inference that the telos is reached. in the late condition, the target picture could only be selected upon hearing the direct object; the predicate could apply to both objects, so the sentence remained ambiguous until the final noun (e.g., a fully colored star vs. a fully colored leaf, cf. figure 1, a-d). an overview of the experimental conditions is provided in figure 1. the audios were recorded by a female native speaker of italian, and manipulated using praat (boersma, 2001), so that the introduction of the sentence (guarda in quale foto ha ‘look in which picture he/she has’) was always the same. the auxiliary ha always started 2400 ms after the start of the trial, and the mean onset of the direct object (i.e., the article preceding the final noun) was at 3668 ms. the mean length of the experimental sentences was 4890 ms. we used a latin square design with three lists of 21 items (seven per condition). participants were assigned to one of the three lists. items were presented in a randomized order. figure 1: target (a) and competitor images in each of the three conditions. 2.3 procedure. children were tested individually in a quiet room within the school, using a portable tobii pro x3-120 eye-tracker which captured their gaze at 120 hz. participants were seated between 60 and 70 cm from the display. calibration took place after a short familiarization phase, consisting of one example and three practice items. during the experimental phase, participants listened to sentences through headphones while their eye movements were recorded. at the end of each sentence, a question mark appeared on the screen, and children could give their offline response by clicking on the mouse to select the correct picture. after that, a fixation cross appeared, ensuring that children were looking at the center of the screen before moving on to the next trial. participants did not receive any feedback about their performance during the experiment. 2.4 analysis. we performed a track loss analysis on the eye-tracking data during the experimental sentences. trials which had more than 35% data loss were removed from the analysis. as a result, 49 trials were removed. moreover, we excluded trials with inaccurate offline responses, which left us with a total of 540 remaining trials for the analysis of the eye gaze data. we used the eyetrackingr (dink & ferguson, 2015) and ggplot2 (wickham, 2016) packages in r (r core team, 2019) to visualize the eye gaze pattern. the data were then analyzed with a competitor in the three conditions target early-lexical early-aspect late a. b. c. d. proceedings of elm 1: 071-077, 2021 jasmijn e. bosch, mathilde chailleux and francesca foppolo: incremental processing of telicity in italian children. 74 https://doi.org/10.3765/elm https://www.elm-conference.net/ generalized linear mixed effect models in r, using the glmer function of the lme4 package (bates, maechler, bolker & walker, 2015). sentences were divided in three time windows: the introduction (guarda in quale foto ‘look in which picture’), the verb (e.g., ha sbucciato ‘has peeled’), and the noun phrase (e.g. la mela ‘the apple’). the boundaries of the time windows were shifted by 200 ms, to take into account the time that is required to plan and execute a saccadic eye movement (altmann, 2011). the statistical analysis tested whether the likelihood of looking at the target (versus competitor) during the noun phrase depended on the experimental condition (early aspect versus early lexical versus late). when comparing early aspect against the baseline condition, late was coded as -.5 and early aspect was coded as +.5. when comparing the two early conditions, early aspect was coded as -.5 and early lexical was coded as +.5. the model also included random intercepts for item and subject. 3. results. the analysis of the offline responses showed a high overall accuracy on the task; 96.7% of the trials were answered correctly. in the analysis of the eye-tracking data, we only focused on accurate trials only. the time course pattern of the eye gaze data is shown in figure 2. figure 2: time course of the proportions of looks toward the target (versus competitor) in the three conditions. the first vertical line indicates the onset of the auxiliary verb; the second vertical line indicates the average onset of the article preceding the final noun. the dotted horizontal line represents chance performance. as can be seen from this plot, in the early lexical condition, participants started directing their gaze toward the target picture during the second time window, which suggests that they immediately integrated the lexical meaning of the verb while processing the sentence. in contrast, in the early proceedings of elm 1: 071-077, 2021 jasmijn e. bosch, mathilde chailleux and francesca foppolo: incremental processing of telicity in italian children. 75 https://doi.org/10.3765/elm https://www.elm-conference.net/ aspect condition and the late condition, we only observe a shift toward the target picture during the noun phrase. nevertheless, participants appear to be faster in the early aspect condition than in the late condition, suggesting that participants were able to make rapid use of the aspectual information on the verb during online sentence processing. in the statistical analysis we focused on the odds of looking at the target during the direct object time window in the three conditions. the summary of the model output is provided in table 1. fixed factor est. odds ratio 95% ci p condition (early aspect vs early lexical) condition (early aspect vs late) 2.06 1.83 1.94 .. 2.18 1.74 .. 1.94 <.001 <.001 table 1: output of the generalized linear mixed model defined as: looks to target (yes or no) ~ condition + (1 | item) + (1 | subject) these results show that participants were significantly more likely to look at the target in the early lexical condition than in the early aspect condition, but they were also significantly more likely to look at the target in the early aspect condition than in the late condition. this confirms the observation that children in this study were sensitive to grammatical aspect, although they were significantly faster when they could rely on simple lexical semantics. 4. discussion. in this paper we presented a visual world eye-tracking experiment conducted with italian children on the incremental processing of telic predicates. previous results with mandarin speaking children and adults show that listeners can distinguish between perfective and imperfective morphemes, providing evidence of a rapid integration of aspectual cues during sentence processing. additionally, previous results with italian adults on similar materials show rapid integration of the completion inference during sentence processing. our aim was to extend previous research to address two experimental questions: (1) is the culmination inference derived incrementally by children? (2) if yes, at which stage of the derivation is it computed? to address question (1) we designed an experiment in which listeners could anticipate the target picture by exploiting linguistic cues on the verb; to address question (2) we contrasted two types of anticipatory cues: one related to verb’s lexical semantics and one related to aspectual (perfective) morphemes on the verb. we discuss three main findings. first, our results show that italian children can anticipate upcoming nouns on the basis of the lexical semantics of verbs, as already observed in classic studies with adults. second, children were faster to shift their gaze to the target in the early-aspect than in the late (control) condition. thus, children relied on the aspectual cue on the verb, since semantics alone did not provide enough information to disentangle the two events depicted in the scenario (recall that in this condition the event was the same, shown at different degrees of completion). this finding provides evidence that children can exploit a morphosyntactic cue denoting perfective aspect to derive the culmination inference that the telos has been reached and the event has been completed, and that they do so incrementally. third, we found earlier anticipatory eye-movements to the target in the early-lexical compared to the early-aspect condition. this result suggests that lexical-semantic information provides a faster cue than aspectual information, and that the culmination inference derived in the aspect condition requires some additional time compared to the integration of basic lexical semantics of the verb. this may reflect an additional cost of the derivation of the culmination inference, which proceedings of elm 1: 071-077, 2021 jasmijn e. bosch, mathilde chailleux and francesca foppolo: incremental processing of telicity in italian children. 76 https://doi.org/10.3765/elm https://www.elm-conference.net/ might take more time for children, despite the fact of being derived incrementally. one possible explanation for this delay relates to the process of integrating visual and linguistic cues during sentence processing, which might be particularly challenging in the early-aspect condition. in this case, two steps are required for the identification of the target: (i) the identification of the event, which is triggered by the lexical semantics of the verb, and (ii) the identification of the degree of completion of the event, which is triggered by morphosyntax (i.e., the combination of auxiliary and past participle). while only step (i) is required in the early-lexical condition, step (ii) is also necessary to anticipate the target in the early-aspect condition. we speculate that this additional step might explain why, at least in children, the effect of anticipation shows up later in a condition that requires the additional integration of aspectual information, in comparison to a condition in which relying on lexical semantics suffices. future research should investigate this issue further. references altmann, gerry t. & kamide, yuki. 1999. incremental interpretation at verbs: restricting the domain of subsequent reference. cognition, 73(3). 247-264. altmann, gerry t. & kamide, yuki. 2007. the real-time mediation of visual attention by language and world knowledge: linking anticipatory (and other) eye movements to linguistic processing. journal of memory and language, 57(4). 502-518. altmann, gerry t. 2011. the mediation of eye movements by spoken language. in the oxford handbook of eye movements. bates, douglas, maechler, martin, bolker, ben & walker, steve. 2015. fitting linear mixed effects models using lme4. journal of statistical software, 67(1). 1-48. boersma, paul. 2001. praat, a system for doing phonetics by computer. glot international 5:9/10. 341-345. dink, jacob w. & ferguson, brock. 2015. eyetrackingr: an r library for eye-tracking data analysis. online: www. eyetracking-r. com. foppolo, francesca, panzeri, francesca, greco, ciro, & carminati, maria n. 2016. the incremental processing of accomplishment predicates. talk at the workshop events in language & cognition, 29th annual cuny conference on human sentence processing. university of florida, gainesville, florida. pickering, martin j., mcelree, brian, frisson, steven, chen, lillian & traxler, matthew j. 2006. underspecification and aspectual coercion. discourse processes, 42(2). 131-155. proctor, andrea s., dickey, michael w. & rips, lance j. (2004). the time-course and cost of telicity inferences. in proceedings of the cognitive science society, 26 (26). 1107-1112. psychology software tools, inc. [e-prime 3.0]. 2020. r core team. 2019. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. townsend, david j. 2013. aspectual coercion in eye movements. journal of psycholinguistic research, 42(3). 281-306. zhou, peng, crain, steven, & zhan, likan. 2014. grammatical aspect and event recognition in children’s online sentence comprehension. cognition, 133(1). 262-276. proceedings of elm 1: 071-077, 2021 jasmijn e. bosch, mathilde chailleux and francesca foppolo: incremental processing of telicity in italian children. 77 https://doi.org/10.3765/elm https://www.elm-conference.net/ listeners use descriptive contrast to disambiguate novel referents and make inferences about novel categories claire bergey* & daniel yurovsky abstract. in the face of unfamiliar language or objects, description is one cue people can use to learn about both. beyond narrowing potential referents to those that match a descriptor, listeners could infer that a described object is one that contrasts with other relevant objects of the same type (e.g., “the tall cup” contrasts with another, shorter cup). this contrast may be in relation to other objects present in the environment or to the referent’s category. in two experiments, we investigate whether listeners use descriptive contrast to resolve reference and make inferences about novel referents’ categories. while participants use size adjectives contrastively to guide novel referent choice, they do not reliably do so using color adjectives (experiment 1). their contrastive inferences go beyond the current referential context: participants use description to infer that a novel object is atypical of its category (experiment 2). overall, people are able to use descriptive contrast to resolve reference and make inferences about a novel object’s category, allowing them to infer new word meanings and learn about new categories’ feature distributions. keywords. pragmatics; reference; concept learning; word learning; description 1. introduction. suppose a friend asked you to “pass the tall dax.” you might look around the room for two similar things that vary in height, and hand the taller one to them. this is how people to respond to adjectives like “tall” with known objects—they preferentially consider candidate referents that have short competitors as soon as they hear “tall” (sedivy et al., 1999). if, on the other hand, there were no objects that varied only in their size, you might infer something different—most daxes must be shorter than the one your friend wants, since people tend to mention atypical features more than typical ones (mitchell et al., 2013; rubio-fernández, 2016). from the indirect information in your friend’s utterance, you could in principle learn either the meaning of a new word, the typical size of a new category, or both. but would you be likely to in practice? in a set of two experiments, we tested whether people use adjectives like “small” and “red” contrastively to determine the meaning of the novel word they heard, whether these adjectives lead people to infer the typical color or size of the described object’s category, and whether these two processes interact. studies using familiar objects show that people use adjective description contrastively to guide their identification of the referent (sedivy et al., 1999; sedivy, 2003). in one such task, four objects appeared on a screen: a target (e.g., a tall cup), a contrastive pair (e.g., a short cup), a competitor that shares the target’s feature but not category (e.g., a tall pitcher), and an irrelevant distractor (e.g., a key). participants then heard a referring expression: “pick up the tall cup.” participants looked more quickly to the correct object when the utterance referred to an object with a same-category contrastive pair (tall cup vs. short cup) than when it referred to an object without a contrastive pair (e.g., when there was no short cup in the display). these results suggest that listeners expect speakers to use description when they are distinguishing between potential referents of the same type, and use this inference to rapidly allocate their attention to the target object * authors: claire bergey, university of chicago (cbergey@uchicago.edu) & daniel yurovsky, carnegie mellon university (yurovsky@cmu.edu). proceedings of elm 1: 039-046, 2021 c©2021 claire bergey and daniel yurovsky published by the lsa with permission of the author(s) under a cc by license. 39 https://doi.org/10.3765/elm https://www.elm-conference.net/ as an utterance progresses. this principle does not apply equally across adjective types, however: color adjectives seem to hold less contrastive weight (sedivy, 2003), perhaps because color adjectives are often used redundantly in english (pechmann, 1989). these experiments demonstrate that listeners use contrast among familiar referents to guide their attention allocation, though not their explicit referent choice, which occurs after the noun disambiguates the object. beyond contrasting a referent with other objects in the present environment, description may draw a contrast between a referent and its category. in production studies, participants tend to describe atypical features more than they describe typical ones (mitchell, reiter, & deemter, 2013; rubio-fernández, 2016; westerbeek, koolen, & maes, 2015). for instance, they almost always include a color descriptor when referring to a blue banana, but not when referring to a yellow one. people therefore use knowledge of contrast with present objects and with an object’s category to inform their production and comprehension of adjectives. can they turn this process around, using the principle of descriptive contrast to learn about novel objects and categories in the world? in this paper, we present a series of experiments to test whether and how listeners make inferences about novel referents using descriptive contrast. first, we examine whether listeners use descriptive contrast to resolve referential ambiguity. in a reference game, participants see groups of novel objects and hear a referring expression asking them to pick one, e.g., “find the small toma.” if participants interpret description contrastively, they should infer that the description was necessary to identify the referent—that the small toma contrasts with some differentlysized toma on the screen. using this contrastive inference, they can resolve referential ambiguity, choosing a small object with a similar larger counterpart rather than a small object with no similar counterpart nearby. second, we test whether listeners use descriptive contrast to make inferences about a novel object’s category. participants are presented with two interlocutors who exchange objects using referring expressions, such as “pass me the blue toma.” if participants interpret description as contrasting with an object’s category, they should infer that in general, few tomas are blue. further, these two inferences may trade off: if the objects in the scene necessitate adjective use to identify the referent uniquely (e.g., there are blue and red tomas in the scene), participants may be less likely to attribute the adjective to a contrast with the category. in order to determine whether people can use contrastive inferences to disambiguate and learn about referents, and how those inferences are affected by adjective type, we use reference games with novel objects. novel objects provide both a useful experimental tool and an especially interesting testing ground for contrastive inferences. these objects have unknown names and feature distributions, creating the ambiguity that is necessary to test referential disambiguation and category learning. but the ability to disambiguate novel referents, or to establish reference with incomplete information, is also the broader problem of learning about the world. we know that distributional information in the world affects people’s pragmatic use and interpretation of description in familiar contexts (sedivy, 2003; westerbeek et al., 2015). here, we ask: can people use pragmatic inferences from description to learn about unfamiliar things in the world? 2. experiment 1. in experiment 1, we tested whether adult participants use adjective contrast to select an ambiguously mentioned novel referent. in a referential disambiguation task, we presented participants with arrays of novel fruit objects (figure 1). on critical trials, participants saw a target object, a lure object that shared the target’s contrast feature but not its shape, and a contrastive pair that shared the target’s shape but not its contrast feature. participants heard a proceedings of elm 1: 039-046, 2021 claire bergey and daniel yurovsky: listeners use descriptive contrast to disambiguate novel referents and make inferences about novel categories. 40 https://doi.org/10.3765/elm https://www.elm-conference.net/ referring expression denoting the feature, e.g., “find the [blue/big] dax.” for the target object, use of the adjective was necessary to disambiguate it from the distractor with the same shape but a different color or size; for the lure, the adjective would be superfluous description. if participants use contrastive inference to choose novel referents, they should choose the target object. however, we do not expect listeners to treat color and size equally. because color is described more superfluously in english than size, we expect size to hold more contrastive weight, encouraging a more consistent contrastive inference. 2.1. method. we recruited 300 participants through amazon mechanical turk. half of the participants were assigned to a condition in which the critical feature was color (stimuli contrasted on color), and the other half were assigned to a condition in which the critical feature was size. stimulus displays were arrays of three novel fruit objects. fruits were chosen randomly at each trial from 25 fruit kinds. ten of the 25 fruit drawings were adapted and redrawn from kanwisher, woods, iacoboni, and mazziotta (1997); we designed the remaining 15 fruit kinds. each fruit kind has an instance in each of four colors (red, blue, green, or purple) and two sizes (big or small). particular target colors were assigned randomly at each trial and particular target sizes were counterbalanced across display types. there were two display types: unique target displays and contrastive displays. unique target displays contained a target object that had a unique shape and was unique on the trial’s critical feature (color or size), and two distractor objects that matched each other’s (but not the target’s) shape and critical feature. contrastive displays contained a target, its contrastive pair (matched the target’s shape but not its critical feature), and a lure (matched the target’s critical feature but not its shape). the positions of the target and distractor items were randomized within a triad configuration. participants were told they would play a game in which they would search for strange alien fruits. each participant saw eight trials. half of the trials were unique target displays and half were contrastive displays. crossed with display type, half of trials had audio instructions that described the critical feature of the target (e.g., “find the [blue/big] dax”), and half of trials had audio instructions with no adjective description (e.g., “find the dax”). participants clicked on the objects to respond. a name was randomly chosen at each trial from a list of eight nonce names: blicket, wug, toma, gade, sprock, koba, zorp, and lomet. after completing the study, participants were asked to select which of a set of alien words they had heard previously during the study. four were words they had heard, and four were novfigure 1. experiment 1 stimuli. on the left: an example of a contrastive trial in which the critical feature is size. here, the participant would hear the instruction “find the small dax.” on the right: an example of a contrastive trial in which the critical feature feature is color. here, the participant would hear the instruction “find the red dax.” in both cases, the target is the top object. proceedings of elm 1: 039-046, 2021 claire bergey and daniel yurovsky: listeners use descriptive contrast to disambiguate novel referents and make inferences about novel categories. 41 https://doi.org/10.3765/elm https://www.elm-conference.net/ el lure words. participants were dropped from further analysis if they did not respond to at least 6 of these 8 correctly (above chance performance as indicated by a one-tailed binomial test at the p = .05 level) or if they missed any of four color perception check trials (resulting n = 163). 2.2. results. we first confirmed that participants understood the task by analyzing performance on unique target trials, in which the target had no competitors with the same shape or critical feature (color or size). we asked whether participants chose the target more often than expected by chance (33%) by fitting a mixed effects logistic regression with an intercept term, a random effect of subject, and an offset of logit(1/3) to set chance probability to the correct level. the intercept term was reliably different from zero for both color (β = 6.64, t = 4.10, p < 0.001) and size (β = 2.25, t = 6.91, p < 0.001). in addition, participants were more likely to select the target when an adjective was provided in the audio instruction in both conditions. we confirmed this effect statistically by fitting a mixed effects logistic regression predicting target selection from condition, adjective use, and their interaction with random effects of participants. adjective type (color vs. size) was not statistically related to target choice (β = -0.48, t = -1.10, p = 0.27), and adjective description in the utterance increased target choice (β = 3.85, t = 3.52, p < 0.001). participants had a general tendency to choose the target in unique target trials, which was strengthened if the audio instruction contained the relevant adjective. our key test was whether participants would choose the target object on contrastive trials in which description was given, reflecting use of a contrastive inference to choose a novel referent (figure 2). to test this, we compared participants’ rate of choosing the target to their rate of choosing the lure, which shares the relevant feature with the target. across all contrast trials, use of an adjective shifted participants toward choosing the target rather than the lure (β = 2.07, t = 6.24, p < 0.001). when size was specified, participants chose the target significantly more often than the lure (β = 0.86, t = 4.41, p < 0.001). however, when color was specified, participants did not choose the target significantly more often than the lure (β = 0.15, t = 0.75, p = 0.45). among contrast trials in which an adjective was not given, participants dispreferred the target, instead choosing the lure object which matched the target’s feature but had a unique shape (β = -2.65, t = -5.44, p < 0.001). participants’ choice of the target over the lure in the size condition was therefore not due to a prior preference for the target in contrast displays, but relied on contrastive interpretation of the adjective. figure 2. experiment 1 contrast trial results. proportion of times that participants chose the target and lure items as a function of adjective condition (color vs. size) and whether an adjective was provided in the utterance. error bars indicate 95% confidence intervals. proceedings of elm 1: 039-046, 2021 claire bergey and daniel yurovsky: listeners use descriptive contrast to disambiguate novel referents and make inferences about novel categories. 42 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.3. discussion. when faced with unfamiliar objects referred to by unfamiliar names, people must resolve ambiguity to understand their conversational partner and learn more about words and their referents in the world. in experiment 1, we tested whether people could use contrastive inferences to resolve ambiguous reference to novel objects. we found that participants have a general tendency to choose objects that are unique in shape when reference is ambiguous. however, when people hear an utterance with description (e.g., “blue toma”, “small toma”), they shift away from choosing unique objects and toward choosing objects that have a similar contrasting counterpart. furthermore, use of size adjectives—but not color adjectives—prompts people to choose the target object with a contrasting counterpart significantly more often than the unique lure object. thus, we found that people are able to use contrastive inferences about size to successfully resolve which unfamiliar object an unfamiliar word refers to. 3. experiment 2. in experiment 1, we examined whether people would interpret description as implying contrast with other present objects. however, as noted, description can imply contrast with sets other than the set of currently available referents. one of these alternative sets is the referent’s category. work by mitchell et al. (2013) and westerbeek et al. (2015) demonstrates that speakers use more description when referring to objects with atypical features (e.g., a yellow tomato) than typical ones (e.g., a red tomato). this selective marking of atypical objects potentially supplies useful information to listeners: they have the opportunity to not only learn about the object at hand, but also about the typical features of its category. in the following experiment, we test whether people use this type of contrast to make inferences about a novel category’s feature distribution. 3.1. method. two hundred and forty participants were recruited from amazon mechanical turk. 120 participants were assigned to a condition in which the critical feature was color (red, blue, purple, or green), and 120 participants were assigned to a condition in which the critical feature was size (small or big). stimulus displays showed two alien interlocutors, one on the left side (alien a) and one on the right side (alien b) of the screen, each with two novel fruit objects beneath them (figure 3). alien a, in a speech bubble, asked alien b for one of its fruits (e.g., “hey, pass me the big toma”). alien b replied, “here you go!” and the referent disappeared from alien b’s side and reappeared on alien a’s side. figure 3. experiment 2 stimuli. in the above example, the critical feature is size and the object context is a within-category contrast: the alien on the right has two same-shaped objects that differ in size. proceedings of elm 1: 039-046, 2021 claire bergey and daniel yurovsky: listeners use descriptive contrast to disambiguate novel referents and make inferences about novel categories. 43 https://doi.org/10.3765/elm https://www.elm-conference.net/ two factors, presence of the critical adjective in the referring expression and object context, were fully crossed within subjects. object context had three levels: within-category contrast, between-category contrast, and same feature. in the within-category contrast condition, alien b possessed the target object and another object of the same shape, but with a different value of the critical feature (color or size). in the between-category contrast condition, alien b possessed the target object and another object of a different shape, and with a different value of the critical feature. in the same feature condition, alien b possessed the target object and another object of a different shape but with the same value of the critical feature as the target. thus, in the withincategory contrast condition, the descriptor is necessary to distinguish the referent; in the between-category contrast condition it is unnecessary but potentially helpful; and in the same feature condition it is unnecessary and unhelpful (see example stimuli in figure 4). all object contexts equated direct observation of the target object category’s feature distribution: participants saw the target object and one other object with the target’s shape and a different critical feature value. we manipulated the critical feature type (color or size) between subjects. participants performed six trials. after each exchange between the alien interlocutors, they made a judgment about the prevalence of the target’s critical feature in the target object’s category. for instance, after seeing a red blicket being exchanged, participants would be asked, “on this planet, what percentage of blickets do you think are red?” and answer on a sliding scale between zero and 100. in the size condition, participants were asked, “on this planet, what percentage of blickets do you think are the size shown below?” with an image of the target object they just saw available on the screen. after completing the study, participants performed the same novel word attention check described in experiment 1, and were excluded if they did not respond to at least 6 out of 8 words correctly (resulting n = 193). figure 4. experiment 2 prevalence judgments. participants consistently judged the target object as less typical of its category when the referent was described with an adjective (e.g., “pass me the purple toma”) than when it was not (e.g., “pass me the toma”). proceedings of elm 1: 039-046, 2021 claire bergey and daniel yurovsky: listeners use descriptive contrast to disambiguate novel referents and make inferences about novel categories. 44 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.2. results. we analyzed participants’ judgments of the prevalence of the target object’s critical feature in its category (figure 4). we fit a maximum mixed-effects linear model, including effects of utterance type (adjective or no adjective), context type (within-category contrast, between-category contrast, or same feature), and critical feature (color or size) as well as all interactions between these three factors and random effects of utterance type and context type nested within subject. random effects were removed until the model converged, resulting in a final model with effects of and interactions between condition, adjective, and context type, and a random effect of utterance type nested within subject. the final model revealed a significant effect of utterance type (βadjective = -11.79, t = -3.90, p < 0.001), such that when an adjective was used, participants inferred that the described feature was less prevalent in object’s category. there was no significant effect of critical feature (βsize = 3.36, t = 1.03, p = 0.301), and there was a significant effect of same-feature context relative to within-category contrast context, but no significant effect of between-category relative to within-category contrast contexts (βsame = -5.41, t = -2.25, p = 0.025; βbetween = -3.92, t = -1.62, p = 0.104). there were no significant interactions. that is, participants slightly adjusted their inferences according to the object context, though not in a way that depended on whether an adjective was used in the utterance. however, they robustly inferred that described features were less prevalent in the target’s category than unmentioned features. 4. general discussion. overall, we found that people are able to use descriptive contrast to infer the referent of a novel word and to make inferences about a novel referent’s category. in our first experiment, participants were able to resolve referential ambiguity using a contrastive interpretation of size adjectives, though not reliably with color adjectives. in our second experiment, participants inferred that a described referent was atypical of its category on that feature: hearing “big toma” led them to think that most tomas were not that size. in real life it is often unclear whether description is meant to contrast with present objects or imply atypicality. in our second experiment, participants adjusted their typicality judgments slightly based on the object context: when the object context was two objects of different kinds with the same critical feature, they judged the target object as slightly more atypical. however, participants did not significantly adjust their prevalence judgments based on the interaction of adjective use and object context— that is, they did not adjust their inferences about typicality based on how redundant description was in context. further, contexts in which description was necessary to identify the referent did not preempt inferences of atypicality. the relative robustness of contrastive inferences about typicality across contexts and adjective types compared to contrastive inferences among present referents raises questions about the relative importance of these two kinds of contrast in language understanding. most prior work has focused on contrast with present referents as the main phenomenon of interest, with object typicality as a modulating factor; our results emphasize the role of contrast with an object’s category, particularly when ambiguity is at play. future work will explore whether people make subtle trade-offs between contrast with present referents and with the referent’s category, and why color–size asymmetries seem to differ between these two types of contrast. further, use of adjectives has been shown to allow children to make contrastive inferences among familiar present objects (huang & snedeker, 2008) and, when paired with contrastive cues such as prosody, about novel object typicality (horowitz & frank, 2016); future work will explore whether adjective contrast alone is a viable learning tool in early childhood. contrastive inferences allow people to learn the meanings of new words and the typical features of new categories, pointing to a broader potential role of pragmatic inference in learning about the world. proceedings of elm 1: 039-046, 2021 claire bergey and daniel yurovsky: listeners use descriptive contrast to disambiguate novel referents and make inferences about novel categories. 45 https://doi.org/10.3765/elm https://www.elm-conference.net/ references huang, y. t., & snedeker, j. (2008). use of referential context in children’s language processing. proceedings of the 30th annual meeting of the cognitive science society. horowitz, a. c., & frank, m. c. (2016). children’s pragmatic inferences as a route for learning about the world. child development, 87(3), 807–819. kanwisher, n., woods, r. p., iacoboni, m., & mazziotta, j. c. (1997). a locus in human extrastriate cortex for visual shape analysis. journal of cognitive neuroscience, 9(1), 133–142. mitchell, m., reiter, e., & deemter, k. van. (2013). typicality and object reference. proceedings of the 35th annual meeting of the cognitive science society. pechmann, t. (1989). incremental speech production and referential overspecification. linguistics, 27(1), 89–110. rubio-fernández, p. (2016). how redundant are redundant color adjectives? an efficiencybased analysis of color overspecification. frontiers in psychology, 7. sedivy, j. c. (2003). pragmatic versus form-based accounts of referential contrast: evidence for effects of informativity expectations. journal of psycholinguistic research, 32(1), 3– 23. sedivy, j. c., k. tanenhaus, m., chambers, c. g., & carlson, g. n. (1999). achieving incremental semantic interpretation through contextual representation. cognition, 71 (2), 109– 147. westerbeek, h., koolen, r., & maes, a. (2015). stored object knowledge and the production of referring expressions: the case of color typicality. frontiers in psychology, 6. proceedings of elm 1: 039-046, 2021 claire bergey and daniel yurovsky: listeners use descriptive contrast to disambiguate novel referents and make inferences about novel categories. 46 https://doi.org/10.3765/elm https://www.elm-conference.net/ the investigation of quantity implicatures during typical development: a systematic review anna teresa porrini & luca surian* abstract. the present work is a systematic review of the acquisition of quantity implicatures in typically-developing children. the references were selected through the prisma method. the criteria for eligibility were that the articles should be peerreviewed, published articles written in english, containing empirical data on the comprehension of quantity implicatures in first language acquisition during typical development. the aim of this review is three-fold. first, to provide a picture of what empirical data tell us about the acquisition of quantity implicatures, based on both lexical and ad-hoc scales, potentially contributing to theoretical accounts of the phenomenon. second, to analyze the methodologies that have been used to test children and their adequacy. lastly, to evaluate whether systematic review is an accurate analysis method for this type of varied and often complicated data. the results suggest that children improve in implicature derivation with age, especially with lexical scales, and that action-based tasks not based on meta-linguistic evaluations might be better suited to test these inferences, especially as opposed to truth value judgment tasks. the fact that the systematic analysis confirms previously individuated trends in the acquisition of implicatures confirms that this is in fact a useful methodology to analyze the data, despite some limitations. keywords. developmental pragmatics; implicatures; quantity maxim; systematic review. 1. introduction. quantity implicatures are enrichments on the meaning of an utterance derived by appealing to the gricean maxim of quantity (grice 1975), which states that, when in conversation, one should make their contribution as informative as is required for the current purposes of the exchange, and not more. quantity implicatures arise by the use of linguistic items often called scalar items, which can be situated within lexical scales – such as the scale but also the and scales, among others – or by the use of contextually dependent, ad-hoc expressions. the following sentences are examples of both types of implicature respectively: (1) john ate some of the cookies. → john ate some but not all of the cookies. (2) in a context with two shirts, one with polka dots and the other with polka dots and stripes. give me the shirt with polka dots. → give me the shirt with polka dots and no stripes. from this section on, for the sake of clarity, we will refer to the former as lexical quantity implicatures and to the latter as ad-hoc quantity implicatures. * anna teresa porrini, università degli studi di trento (annateresa.porrini@unitn.it) & luca surian, università degli studi di trento (luca.surian@unitn.it). proceedings of elm 2: 219-228, 2023 c©2023 anna teresa porrini and luca surian published by the lsa with permission of the author(s) under a cc by license. 219 https://doi.org/10.3765/elm https://www.elm-conference.net/ data on acquisition of quantity implicatures suggest that children may have difficulties in deriving them correctly and often interpret as meaning (e.g., guasti et al. 2005, huang & snedeker 2009, noveck 2001, papafragou & musolino 2003, sullivan et al. 2019). this seems to be in contrast with the fact that children are very capable from younger ages when it comes to other pragmatic abilities (condry & spelke 2008, matthews et al. 2012, tomasello 2003), and begs the question of whether their difficulties with quantity implicatures may be due not to general pragmatic language delay, but to some other factors. for instance, the data underline a distinction between lexical and ad-hoc scales, with the former being significantly more difficult than the latter (foppolo et al. 2020, horowitz et al. 2017, kampa & papafragou 2019, stiller et al. 2015, wilson & katsos 2021, yoon & frank 2019, zhao et al. 2021). in a way, this seems to suggest that the difficulty children have with implicatures lays in their lexical knowledge. at the same time, however, the only corpus study available so far suggests that in production children are competent in their use of lexical scales from a very young age (eiteljeorge et al. 2018). as a matter of fact, several hypotheses have been put forward to explain children’s delay in acquiring lexical quantity implicatures (barner et al. 2011, foppolo et al. 2012, katsos & bishop 2011, pouscoulous et al. 2007, reinhart 2004, skordos & papafragou 2016 among others). research to disentangle this issue is still ongoing, and experimental data on the acquisition of quantity implicatures are abundant. there is, however, great variety within the available data: children have been tested in different languages, at different ages and with different tasks, and the phenomenon of quantity implicature has been presented using different scales and in cooccurrence with other linguistic or cognitive factors. the various layers of complication present in the literature make it difficult to compare studies directly. still, the considerable number of studies conducted allows for an analysis of available data in the form of a systematic review, which could give the possibility of a more objective viewpoint on the available data as a whole. the methodology is, however, not without its limitations, as it does not allow for analyses of the details of each study and of some of the relevant differences between different experiments, tasks and items. the aim of this review is therefore not only to shed light on the phenomenon of quantity implicatures during typical development, and how they have been tested so far, but also to validate systematic reviewing as a methodology for analyzing available data. 2. methodology. through a synthetic search and evaluation of multiple studies, we concentrated on quantity implicatures in an attempt to describe the data collectively. the references for this review were selected through the prisma method, and initially the search was extended to any type of implicature.1 at first, a search through keywords was performed on three databases: scopus, web of science and apa psycinfo. then, the data were screened in a four stage process, in order to only include articles that fit our eligibility criteria. after collecting the articles and extracting the data, we decided to concentrate only on quantity implicatures. the main reason for this was that the data were much richer for this type of implicature and allowed for more interesting comparisons. there were various eligibility criteria selected for this search: first, the selected references needed to be peer-reviewed, published articles written in english after the year 2000. second, they should contain empirical quantitative data on the comprehension of implicatures in first language acquisition during typical development. moreover, there needed to be a clear classifica 1 prisma flowcharts, as well as the final dataset and r script, are available on osf: https://osf.io/g4edn/. proceedings of elm 2: 219-228, 2023 anna teresa porrini and luca surian: the investigation of quantity implicatures during typical development. 220 https://doi.org/10.3765/elm https://www.elm-conference.net/ tion of what type of implicature was being tested and in what way, with examples. finally, the authors needed to have performed a replicable statistical analysis on the data and there needed to be indication of the age range and mean age of the participants. in order to make the data more easily comparable, one last criterion was that the articles should all present their results in terms of percentage of success in implicature derivation (or a measure that could be converted to this). 3. results. 3.1. collected database. in the end, 44 papers were deemed eligible for the analysis, all published between the years 2001 and 2021.2 within these references, a total of 158 different findings in terms of percentage of success was obtained across the different experiments, implicature types, tasks and groups tested within the 44 references. the minimum age tested was 2 years old and the maximum age tested was 13 years and 4 months old. information on how many findings were found for each age group can be found in table 1. mean age in years findings per age group 2 2 3 11 4 42 5 54 6 9 7 21 8 3 9 5 10 8 11 3 table 1: distribution of findings by age the experiments were run in eight different languages: dutch, english, french, greek, italian, japanese, mandarin chinese and spanish. the results are generalizable beyond the scope of just one language, as there is no detectable difference in percentage of success among the eight languages. in fact, while a kruskal-wallis rank sum test reports a chi-square of 17.102 and a pvalue of 0.017, which shows a significant effect of language on performance, a subsequent dunn test reveals that there is no statistically significant difference between any two languages. six different task types were used to test quantity implicatures within the dataset. a summary of the tasks used and how many findings were collected with each can be seen below in table 2. the most frequently used tasks were the truth value judgment task (tvjt), the felicity judgment task (felj) and the referent selection task (refs). in tvjt experiments, participants are asked to make a binary choice regarding the truthfulness (or correctness) of an uttered sentence, while in felj they are simply asked to evaluate whether the speaker had “said something well”, and therefore to make a judgment on how felicitous the sentence was, rather than true. the difference between the two tasks is not always unequivocal, but experimenters tend to specify which sentences they used to ask for judgment, which made the categorization of the two during data collection for the present review less arbitrary. referent selection task 2 the data collection was performed in august 2021. proceedings of elm 2: 219-228, 2023 anna teresa porrini and luca surian: the investigation of quantity implicatures during typical development. 221 https://doi.org/10.3765/elm https://www.elm-conference.net/ (refs), on the other hand, is in this instance an umbrella term for any task that required participants to pick a referent for a determinate sentence or utterance, be it a picture, a person or an object. as for the less used tasks, action based tasks (actb) are those in which children are asked to perform an action after hearing a sentence, instead of being asked for judgment. in communicative context assessment (cca), children are asked to give a judgment on an event that took place after a sentence (an instruction) was uttered, instead of after the sentence itself, as is the case for tvjt and felj. the last task is the speaker selection task (spes), in which participants need to select which of two speakers uttered a determinate sentence, and it can be for instance that one of the speakers has full knowledge of what happened, while the other does not. within these six task types, four possible types of output variable types were presented to participants: binary, ternary, quaternary or performative. task findings per task action based 10 communicative context assessment 2 felicity judgment 43 referent selection 54 speaker selection 9 truth value judgment 40 table 2: distribution of findings by task 3.2. data analysis. the data analysis was performed using rstudio (r 4.1.0) in different ways: first, a generalized linear model was fitted to analyze the data as a whole, inserting percentage of success as a dependent variable and mean age, task type and implicature type (lexical or adhoc) as independent variables. variable type and language were not selected as factors for the glm because their inclusion did not guarantee a better fit. the results of this analysis can be seen in table 3. the results of the glm suggest that age does have an effect on performance, as well as implicature type and task type. then, non-parametric tests were performed at different stages of the analysis. we will discuss the effects of age, implicature type and task in the following section. estimate std. error z value p-value (intercept) 1.085 1.074 1.010 0.313 mean age 0.017 0.008 2.066 0.039 * lexical scale -1.115 0.475 -2.346 0.019 * cca -0.255 1.688 -0.151 0.880 felj -1.047 0.839 -1.248 0.212 refs -0.982 0.848 -1.158 0.247 spes -0.924 1.059 -0.872 0.383 tvjt -1.538 0.841 -1.828 0.068 . table 3: results of the generalised linear model (with ad-hoc scale and actb as default) proceedings of elm 2: 219-228, 2023 anna teresa porrini and luca surian: the investigation of quantity implicatures during typical development. 222 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4. discussion. 4.1. age. it appears from the analysis that age does, in fact, have an effect on children’s performance with implicature derivation. this should not come as a surprise, since improvements with age were detected in virtually every experiment on the subject. as can be seen in figure 1, the effect is confirmed to be a positive one. figure 1: effect of age 4.2. implicature type. the glm seems to suggest that ad-hoc implicatures are easier to derive as compared to lexical ones, since the lexical implicature type has a significative negative effect in the model. a wilcox rank test confirmed that the difference in percentage of success between the two implicature types is significant (p < 0.001). if we look at the data more closely, however, this difference is less and less detectable as children get older, as seen in figure 2. after dividing the dataset by age, roughly based on figure 2: difference between scales at different age ranges proceedings of elm 2: 219-228, 2023 anna teresa porrini and luca surian: the investigation of quantity implicatures during typical development. 223 https://doi.org/10.3765/elm https://www.elm-conference.net/ whether children could already be in school or not (age 2 to 5;11 on one side, age 6 to 13;4 on the other), two new wilcox rank tests show that in the case of preschool children the difference is indeed significant (p < 0.001), while it is not for older children (p=0.5). another thing worth noticing on the topic of implicature type is that there is a number of horn scales that might be used to test lexical quantity implicatures. in the majority of cases, however, children are tested on the scale. to be more precise, out of 39 papers that experimented using lexical scales, 26 tested only the scale, thus only 13 of the remaining also included the scales and modal or aspectual verbs. a kruskal-wallis rank sum test was performed on the results for lexical quantity implicatures, analyzing the difference in performance based on the scales used. this resulted in a chi-square of 14.872, and a p-value of 0.005, signaling that there is indeed a difference in how successful children are based on the lexical scale used. a subsequent dunn test confirmed that the significance was driven by the difference between the and the and modal verb scales. the scale appears significantly easier than the other two for children (p=0.0123 for modals, p=0.0012 for ). the lack of more data on scales other than , however, makes more accurate analyses on the potential differences difficult to perform. 4.3. task. with regards to task type, the glm shows a marginally significant negative effect of the tvjt. this task has been previously criticized in the literature for its inaccuracy when dealing with pragmatic phenomena (e.g. katsos & bishop 2011). one of the issues of the tvjt is that it requires participants to make judgments about truthfulness, and for this reason it might fail to detect fundamental aspects of pragmatic processing. pragmatic meaning, in fact, is concerned with informativeness, felicity and optimality more than it is with truthfulness itself, which is more a domain of semantics. there is work that provides proof that the tvjt is indeed a good methodology to test implicatures (guasti et al. 2005, foppolo et al. 2020), but the analysis of systematically collected data does not support this view. in fact, it seems that the among the three most used tasks in the dataset, the tvjt is the only one for which age does not predict performance, and children show lower accuracy even at older ages (figure 3 for reference). figure 3: mean success rates and boxplots at different ages for tvjt proceedings of elm 2: 219-228, 2023 anna teresa porrini and luca surian: the investigation of quantity implicatures during typical development. 224 https://doi.org/10.3765/elm https://www.elm-conference.net/ interestingly, an exploratory analysis of the data seems to suggest that some of the less used methodologies might grant children better performance on quantity implicatures. as figure 4 shows, children seem to perform better with action based tasks (actb) and communicative context assessment tasks (cca). these are both tasks that do not require children to make any meta-linguistic judgments, which might be a source of difficulties at younger ages. a kruskalwallis rank sum test was performed on the data to verify whether task had an effect on success with implicature derivation. this test showed that the difference between tasks is indeed significant (chi-square = 13.595, df = 5, p-value = 0.0184). more in-depth analysis through a dunn test revealed that the significant difference that drives this effect is between tvjt and actb, in line with the result of the glm (p= 0.0118). in speaker selection tasks children do not seem to be as good as they are with actb and cca, and this may be due to the fact that it requires more meta-representational abilities such as reasoning on other people’s knowledge. this is a very speculative analysis, however, because the scarceness of data for all three task types hardly allows for any conclusions to be drawn from statistical analysis. 4.4. variable type. output variable type was not included as a factor in the glm because it did not guarantee a better fit of the model with the data, nor did it seem to interact with other factors such as task type or scale. previous literature, however, suggests that in felicity judgment tasks (felj), when children are asked to judge someone’s utterance by giving them a prize, they perform significantly better when they are given the opportunity to express judgment on a ternary scale (small prize, medium prize and big prize) as opposed to a binary choice (prize or no prize). when given a ternary option, children seem to be able to distinguish between an incorrect sentence, a correct sentence and a non-optimal but correct sentence, which is pragmatically infelicitous. an analysis of the available data confirms this, as can be seen from figure 5. children’s performance in felj tasks when the outcome variable is ternary is better than when it is binary. figure 4: difference in performance across tasks proceedings of elm 2: 219-228, 2023 anna teresa porrini and luca surian: the investigation of quantity implicatures during typical development. 225 https://doi.org/10.3765/elm https://www.elm-conference.net/ a wilcoxon rank sum test with continuity correction was performed to confirm this effect, and it resulted in a p-value of 0.0124. this suggests that the difference is indeed significant, although it needs to be specified that within the dataset the instances of binary outcome variables were considerably more compared to the ternary ones. 5. conclusion. what the systematic review as a whole seems to point to is that children improve in deriving quantity implicature with age, which is an expected result. it also confirms that there is a detectable difference between ad-hoc and lexical scales that makes the latter more difficult for children, especially in their preschool years. another important result of this systematic analysis of experiments on quantity implicatures during development regards experimental factors. it seems that the task used to test children may have a considerable impact on how they perform in implicature tasks. what the literature suggests, in this case, is that tasks that do not rely on children’s meta-linguistic or metarepresentative abilities may be better suited for these investigations. the data also points to an inadequacy of tvjt, a task that revolves around truthfulness instead of felicity. although the result is relevant, as it reveals a trend across different experiments throughout the past 20 years, it needs to be pointed out that the methodology of systematic reviewing ignores the presence of potential experimental manipulations that might make the task more or less adequate to test implicatures on children. on another note implicit tasks would probably provide good results as well. these tasks were not included in this review due to their outcome measure being too different from the others, and thus not easily comparable to the rest of the dataset, but based on the reasoning that meta-linguistic and meta-representative judgments exert a significant toll on children, an implicit methodology such as eye-tracking might be an optimal way of testing them. as a conclusion, we would argue that the systematic methodology is probably not sufficient by itself to analyze data in depth, as it does not afford attention to details, features and manipulations within each experiment which might be of considerable interest to researchers. nevertheless, it is an extremely useful and apparently accurate methodology if the aim is to get a broader, comprehensive view of the available data, in which salient factors to be kept in mind can be individuated at a glance. a systematic review can be an excellent starting point for refigure 5: success rates for binary and ternary outcome variables in felj proceedings of elm 2: 219-228, 2023 anna teresa porrini and luca surian: the investigation of quantity implicatures during typical development. 226 https://doi.org/10.3765/elm https://www.elm-conference.net/ search on a topic, as it also shows in what aspects the current literature could be enriched: with regards to quantity implicatures during development, we would suggest that this review recommends the following. it could be interesting to test a wider variety of lexical scales, for one, and to use more implicit tasks or tasks that do not require meta-linguistic abilities. another interesting suggestion might be to try and deviate from binary outcome variables more often. these future directions might be a good step in the direction of getting a more comprehensive picture of what children are really capable of in terms of quantity implicature comprehension. references barner, d., brooks, n., & bale, a. (2011). accessing the unsaid: the role of scalar alternatives in children’s pragmatic inference. cognition 118. 87–96. https://doi.org/10.1016/j.cognition.2010.10.010. condry, k. f., & spelke, e. s. (2008). the development of language and abstract concepts: the case of natural number. journal of experimental psychology: general 137(1). 22–38. https://doi.org/10.1037/0096-3445.137.1.22. eiteljoerge, s. f. v., pouscoulous, n., and lieven, e. v. m. (2018). some pieces are missing: implicature production in children. frontiers in psychology 9:1928. https://doi.org/10.3389/fpsyg.2018.01928. foppolo, f., guasti, m. t., & chierchia, g. (2012). scalar implicatures in child language: give children a chance. language learning and development 8. 365–394. https://doi.org/10.1080/15475441.2011.626386. foppolo, f., mazzaggio, g., panzeri, f., & surian, l. (2020). scalar and ad-hoc pragmatic inferences in children: guess which one is easier. journal of child language 48(2). 350-372. http://dx.doi.org/10.1017/s030500092000032x. grice, h. p. (1975). logic and conversation. in p. cole & j. l. morgan (eds.). syntax and semantics: speech acts (vol. 3, pp. 41–58). new york: academic press. guasti, t. m., chierchia, g., crain, s., foppolo, f., gualmini, a., & meroni, l. (2005). why children and adults sometimes (but not always) compute implicatures. language and cognitive processes 20(5). 667–696. https://doi.org/10.1080/01690960444000250. horowitz, a. c., schneider, r. m., & frank, m. c. (2017). the trouble with quantifiers: exploring children’s deficits in scalar implicature. child development 8(6). e572–e593. https://doi.org/10.1111/cdev.13014. huang, y. t., & snedeker, j. (2009). semantic meaning and pragmatic interpretation in 5-yearolds: evidence from real-time spoken language comprehension. developmental psychology 45(6). 1723–1739. https://doi.org/10.1037/a0016704. kampa, a., & papafragou, a. (2019). four‐year‐olds incorporate speaker knowledge into pragmatic inferences. developmental science 23(3). e12920. https://doi.org/10.1111/desc.12920. katsos, n., & bishop, d. v. (2011). pragmatic tolerance: implications for the acquisition of informativeness and implicature. cognition 120. 67–81. https://doi.org/10.1016/j.cognition.2011.02.015. matthews, d., butcher, j., lieven, e., & tomasello, m. (2012). twoand four-year-olds learn to adapt referring expressions to context: effects of distracters and feedback on referential communication. topics in cognitive science 4(2). 184–2010. https://doi.org/10.111/j.17568765.2012.01181.x. noveck, i. a. (2001). when children are more logical than adults: experimental investigation of scalar implicature. cognition 78. 165-188. https://doi.org/10.1016/s0010-0277(00)00114-1. proceedings of elm 2: 219-228, 2023 anna teresa porrini and luca surian: the investigation of quantity implicatures during typical development. 227 https://doi.org/10.3765/elm https://www.elm-conference.net/ papafragou, a., & musolino, j. (2003). scalar implicatures: experiments at the semantics– pragmatics interface. cognition 86. 253–282. https://doi.org/10.1016/s00100277(02)00179-8. pouscoulous, n., noveck, i., politzer, g., & bastide, a. (2007). processing costs and implicature development. language acquisition 14. 347–375. https://psycnet.apa.org/doi/10.1080/10489220701600457. reinhart, t. (2004). the processing cost of reference set computation: acquisition of stress shift and focus. language acquisition 12. 109–155. https://doi.org/10.1207/s15327817la1202_1. skordos, d., & papafragou, a. (2016). children’s derivation of scalar implicatures: alternatives and relevance. cognition 153. 6–18. https://doi.org/10.1016/j.cognition.2016.04.006. stiller, a., goodman, n. d., & frank, m. c. (2015). ad hoc implicature in preschool children. language, learning and development 11(2). 176–190. https://doi.org/10.1080/15475441.2014.927328. sullivan, j., davidson, k., wade, s., & barner, d. (2019). differentiating scalar implicature from exclusion inferences in language acquisition. journal of child language 46. 733–759. https://doi.org/10.1017/s0305000919000096. tomasello, m. (2003). constructing a language: a usage based theory of language acquisition. cambridge, ma: harvard university press. wilson, e., & katsos, n. (2021). pragmatic, linguistic and cognitive factors in young children's development of quantity, relevance and word learning inferences. journal of child language. 1-28. https://doi.org/10.1017/s0305000921000453. yoon, e. j., & frank, m.c. (2019). the role of salience in young children’s processing of ad hoc implicatures. journal of experimental child psychology 186. 99–116. https://doi.org/10.1016/j.jecp.2019.04.008. zhao, s., ren, j., frank, m. c., & zhou, p. (2021). the development of quantity implicatures in mandarin-speaking children, language learning and development. 343-365. https://doi.org/10.1080/15475441.2021.1886935. proceedings of elm 2: 219-228, 2023 anna teresa porrini and luca surian: the investigation of quantity implicatures during typical development. 228 https://doi.org/10.3765/elm https://www.elm-conference.net/ modeling the role of polysemy in verb categorization elizabeth soper & jean-pierre koenig* abstract. recent work has indicated that static word embeddings can predict human semantic categories (majewska et al. 2021). in this paper, we consider the role of polysemy in semantic categorization, by comparing sense-level embeddings with previously studied static embeddings in their prediction of human-produced categories. we find that the polysemy is crucial for predicting human categories; sense-level embeddings dramatically outperform static embeddings in predicting semantic categories. our findings highlight the role of polysemy in semantic categorization that is exclusively based on linguistic input. keywords. distributional semantics; polysemy; categorization; natural language processing. 1. introduction. as we learn language and learn about the world, we acquire semantic knowledge. this knowledge comes both from perceptual input (what we experience first-hand) as well as linguistic input (what we hear and read about). we organize our knowledge of categories, and the words we use to denote them, based on connections we form between words and the things they denote in the world, as well as connections between words and other words. in this paper, we are interested in exploring two interrelated issues: (1) how much of human categories may come from what we hear or read? and (2) whether partitioning word representations into similar contexts of use (a proxy for fine-grained polysemy) improves the fit of linguistic knowledge to human categories? distributional semantic models are well-suited to help us answer these two questions as they use linguistic co-occurrence statistics from a corpus to create vector representations of word meanings. words with similar meanings, which occur in similar contexts, end up near each other in space, while unrelated words, which occur in very different contexts, end up far apart in space. word embeddings have become popular in recent years, in particular for their ability to accurately predict the similarity between words (landauer & dumais 1997, pereira et al. 2016, devlin et al. 2019). since similarity is a primary criterion for categorization (collins & loftus 1975), word embeddings may also be good at predicting semantic categories. because word embeddings are trained on linguistic input only, with none of the perceptual input that humans receive, they are ideal for studying the unique role of language in semantic categorization. additionally, because current word embeddings keep track of the contexts in which words are found, we can look specifically at how polysemy affects how semantic categorization can derive from linguistic knowledge. as words generally have multiple possible senses, categorization decisions may depend on which which sense of a word is being considered. representing the distinct senses of polysemous words is thus likely to be important to how humans categorize sets of verbs’ denotations and how these categories to denotations are used to approximate categories *the authors gratefully acknowledge the audience at the experiments in linguistic meaning conference, hosted by the university of pennsylvania, for their useful comments and feedback. authors: elizabeth soper, suny at buffalo (esoper@buffalo.edu) & jean-pierre koenig, suny at buffalo (jpkoenig@buffalo.edu). proceedings of elm 2: 278-287, 2023 c©2023 elizabeth soper and jean-pierre koenig published by the lsa with permission of the author(s) under a cc by license. 278 https://doi.org/10.3765/elm https://www.elm-conference.net/ of objects. we compare different types of word embeddings when it comes to predicting human categories, and find that representing individual senses is indeed crucial for predicting semantic categories. 2. background. there are two main classes of word embedding models: traditional static embeddings (landauer & dumais 1997, mikolov et al. 2013, pennington et al. 2014) represent each word type as a unique vector, while more recent contextual models (peters et al. 2018, devlin et al. 2019) generate a unique representation for every instance of a word in context. since both static and contextual embeddings have been shown to model pairwise similarity between words well (pereira et al. 2016, chronis & erk 2020), and since similarity is a primary criterion for categorization, it seems intuitive that word embeddings should predict categorization well. some previous work supports this intuition; word embeddings have excelled at word sense disambiguation (giulianelli et al. 2020, soler & apidianaki 2021, chronis & erk 2020) and topic modeling (sia et al. 2020, aharoni & goldberg 2020), when cast as categorization problems. as mentioned in the introduction, we are interested in the present paper in linguistically-based semantic category induction. instead of grouping individual word tokens into distinct senses, or documents into topics, the goal of semantic categorization is to group unique words into semantically related clusters. this more abstract type of categorization has received less attention in the word embedding literature. in order to understand what word embeddings can tell us about the role of polysemy in deriving semantic categories from linguistic input, it is important to note that there are at least two different approaches to explaining polysemy. one account holds that polysemous words have a single, under-specified meaning, and that context helps disambiguate between a defined set of senses (pustejovsky 1998). static models are analogous to this view of meaning, as they create a single representation which is meant to encompass all uses of a word form. a stronger claim has been made (for example, by elman 2009; see also marvel & koenig 2015) that different senses of a word are not simply reflected in, but actually created by context. proponents of this claim believe that word meaning is fundamentally context-dependent. contextual language models like bert implicitly take this view of word meaning, as they represent each instance of a word in a particular context as a unique embedding. we can then treat classes of contexts as equivalent to senses in these models. because static and contextual models align with different theoretical perspectives on polysemy, word embeddings are well-suited to test these two conceptions of polysemy and their effect on linguistically-based category induction; static embeddings are a proxy for the single-entry view, and contextual embeddings for the radically context-dependent view. previous work using word embeddings to model semantic categorization has treated semantic categorization as categorization of word types, rather than word senses. by using static embeddings only, it implicitly assumed that people use a summary representation to categorize words. while summary or underspecified representations of word meanings is appropriate for many tasks, as pustejovsky (1998) shows, much work in psycholinguistics suggests that when reading words in and out of context, native speakers often favor one sense or the other (see, among others, brocher et al. 2018 for evidence and summary). in using linguistic input to derive (part of) their semantic knowledge, learners therefore most likely rely on the word sense instantiated in the input. as a result, categorization decisions proceedings of elm 2: 278-287, 2023 elizabeth soper and jean-pierre koenig: modeling the role of polysemy in verb categorization. 279 https://doi.org/10.3765/elm https://www.elm-conference.net/ may depend on which sense of a word is being considered and any attempt to determine how much semantic knowledge may derive from linguistic input would be best served by taking polysemy into consideration. if this is the case, we would expect contextual embeddings to predict semantic categories better than static embeddings. interestingly, recent work evaluating different word embedding models on verb categorization suggests just the opposite. majewska et al. (2021) found that contextual models perform poorly compared to older static models when approximating the verb categorization done by participants in their experiments. we argue that this result is due not to the irrelevance of context to categorization, but rather to the way the contextual embeddings were extracted from the model in majewska et al. (2021). although many of the words in majewska et al. (2021)’s ground truth data are polysemous and are assigned to multiple categories by participants, they evaluate models in a one-representation-per-word-form manner. even when evaluating bert, which has been shown to encode sense-specific information (chronis & erk 2020, soler & apidianaki 2021), this information was thrown away, either by feeding words to the model in isolation or by averaging over all contexts. because they use polysemous data to test representations which do not encode sense information, majewska et al. (2021)’s results may not reflect the full potential of contextual architectures to model categorization. in this study, we test the ability of static and sense-specific embeddings to predict humanproduced categories. our ultimate goal is to determine the role linguistic input plays in semantic knowledge; our more modest goal in this paper is to find out whether polysemy matters in modeling verb categorization. our overall finding is that, as we predicted, sense-specific embeddings are much better at predicting human behavior than static embeddings. 3. dataset. we use the phase 1 data from spa-verb (majewska et al. 2021), which contains 825 verbs sorted into 17 broad classes. 10 participants were asked to sort words by clicking and dragging each word into a circle (see figure 1). there were no constraints on the number or size of groups participants could create; they were only asked to sort them according to their meaning. table 1 gives an overview of the classes resulting from this sorting task. 116 verbs belong to more than one class. no words were assigned to more than 3 classes. on average, words are assigned to 1.14 classes. the spa-verb dataset is particularly relevant for our goals because the categories reflect actual human behavior, rather than expert-curated categories. additionally, since words were presented in isolation with no disambiguating context, using this dataset is a strong test of the role of polysemy. if people categorize according to individual senses rather than a summary representation of each word, even when shown words out of context, this would support a context-sensitive view of meaning and suggest that polysemy matters when building semantic knowledge from linguistic input. 4. models. next we describe the models we compared against our baseline of human performance. 4.1. word2vec. the first model we evaluate is a word2vec model trained on part-of-speechtagged data (fares et al. 2017). part-of-speech tagging allows the static model to distinguish between senses which have different parts of speech (e.g. duck noun and duck verb), alproceedings of elm 2: 278-287, 2023 elizabeth soper and jean-pierre koenig: modeling the role of polysemy in verb categorization. 280 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: screen interface for the spa-verb sorting task (majewska et al. 2021) cluster label example verbs movement wander, fly, glide, roam communication persuade, command, tell crime & law beat, abduct, abuse, shoot negative emotion offend, aggravate, enrage positive emotion admire, respect, adore, like cognitive process suppose, assume, realize cooking cook, slice, stew, boil possession belong, obtain, acquire table 1: a sample of the 17 gold classes in spa-verb dataset (labels are given for descriptive purposes only) proceedings of elm 2: 278-287, 2023 elizabeth soper and jean-pierre koenig: modeling the role of polysemy in verb categorization. 281 https://doi.org/10.3765/elm https://www.elm-conference.net/ though senses which have the same part of speech are still conflated into a single vector (e.g. get#acquire and get#understa-nd). skip-gram with negative sampling was used to train the model on gigaword 5th edition (parker et al. 2011), with a context window size of 5 and 300 dimensions. six words from spa-verb (do, have, be, own, slurp, and snooze) were not in the model’s vocabulary (due either to removal of stop words from corpus or rare words not occurring in corpus) and were thus excluded from our analysis. 4.2. bert. we evaluate three methods of extracting embeddings from the contextual bert model: two baseline methods, which create one representation per word form, and a multi-prototype method which generates one representation per word sense. for all methods, we use bert base uncased from huggingface’s transformers package (wolf et al. 2020). 4.2.1. decontextualized (decont). first and most simply, we extract embeddings from bert by feeding each word to the model in isolation. this creates a single, static embedding for each word. this strategy has been used previously as a way to easily extract ‘context-free’ representations from bert (liu et al. 2019, vulić et al. 2020). 4.2.2. aggregated (aggr). next, we create static embeddings from bert by averaging a word’s embeddings across 100 unique contexts. this aggregated approach still reduces a word to a single representation, but has been shown to produce higher quality representations than the decontextualized strategy (bommasani et al. 2020). 4.2.3. multiprototype (mpro). finally, to test whether sense-specific information is important to semantic categorization, we distill token-level bert embeddings into multiple prototype embeddings. we use the method of chronis & erk (2020) to generate representations which correspond to different senses of a word, without collapsing every token into a single representation. multi-prototype embeddings were generated as follows: 1. for each verb in the dataset, we sampled up to 100 sentences from the british national corpus (bnc consortium 2007), excluding non-verbal uses of the target word. a few words in the set occurred in the corpus fewer than 100 times. four words (broil, corrupt, exhale, and misspend) did not occur as verbs at all in the corpus and were excluded from our analysis. the average number of occurrences sampled for a word was 95.6. 2. we extract bert token embeddings for each collected occurrence of a word. for words which bert tokenizes into multiple word pieces, we average over all component pieces. 3. we cluster the token embeddings for each verb. like chronis & erk (2020), we use k-means clustering to group tokens into ‘sense’ clusters. we use the number of verb senses listed in wordnet (miller 1995) to determine the appropriate k for each word. verbs in the dataset had on average 5.9 senses. (min: 1, max: 59, for buzz). 4. after identifying clusters, we take the k cluster centroids for each word. these are the embeddings we evaluate against the spa-verb categorization data. 4.3. random baseline. finally, we generate random vectors and evaluate them in order to establish a baseline for random chance performance. 5. evaluation. to evaluate the performance of each model against a ‘gold standard’ set of humangenerated categories, k-means clustering is used to group verbs into predicted classes. we use the proceedings of elm 2: 278-287, 2023 elizabeth soper and jean-pierre koenig: modeling the role of polysemy in verb categorization. 282 https://doi.org/10.3765/elm https://www.elm-conference.net/ same metrics as majewska et al. (2021): modified purity and weighted class accuracy are combined in an f1 score, calculated as their balanced harmonic mean. modified purity is the mean precision of predicted clusters: mpur = ∑ c∈clust,nprev(c)>1nprev(c) #test verbs (1) where each cluster c from the set of all kclust induced clusters clust is associated with its prevalent gold class, and nprev(c) is the number of verbs in an induced cluster c taking that prevalent class, with all other verbs considered errors. #test verbs is the total number of verbs in the dataset. while modified purity is a measure of precision, weighted class accuracy targets recall: wacc = ∑ c∈goldndom(c) #test verbs (2) where for each class c from the set of gold standard classes gold, we identify the dominant cluster from the set of induced clusters having most verbs in common with c (ndom(c)). because mpro bert has multiple representations for a single word, the same word form may show up more than once within a single cluster. to prevent artificially inflating the recall in evaluating mpro bert, we eliminate duplicates within each cluster before evaluation. 6. results. table 2 shows the results of each embedding type, compared to results reported in majewska et al. (2021). the baseline models (decont. and aggr. bert) perform comparably to previously reported results. part-of-speech-sensitive word2vec model scores about 10 points higher than reported for a similar model architecture without part-of-speech information. mpro bert, by contrast, performs dramatically better than other embeddings, achieving more than double the f1 score of the best previously reported bert results. this suggests that polysemy does play an important role in modeling linguistically-based semantic categorization: human categories that can be derived from linguistic input correspond to semantically similar contexts of use of words. model f1-optimal f1-gold random baseline 0.204 0.161 majewska word2vec 0.355 0.326 majewska best bert 0.340 0.322 pos-tagged word2vec 0.442 0.433 decont. bert 0.309 0.191 aggr. bert 0.398 0.346 mpro bert 0.743 0.687 table 2: average f1 across models on coarse-grained categories. ‘gold’ is for k=17, as in the ground truth. ‘optimal’ is best result for k in the range (5, 50). the benefit of sense-specific embeddings for modeling coarse-grained categoirzation is clear in the example of freeze. in the ground truth data, freeze belongs to just one class, related to cooking (along with words like bake, fry, melt, and thaw). freeze has another figurative sense, meaning to stop or suspend. because the word is polysemous, static embedding clusters struggle to categorize proceedings of elm 2: 278-287, 2023 elizabeth soper and jean-pierre koenig: modeling the role of polysemy in verb categorization. 283 https://doi.org/10.3765/elm https://www.elm-conference.net/ it appropriately. in the aggregated bert clusters, freeze appears in a cluster predominated by verbs related to violence (whip, shoot, choke, crush, smash). decontextualized bert puts freeze in a heterogeneous cluster with a few cooking words (melt, stew, fry) but also many seemingly unrelated words (knit, greet, disturb, wander). it appears that the different senses of the word skew its static representation and prevent accurate classification. mpro bert, by contrast, puts freeze in two clusters: one related to cooking (as in the ground truth) and another cluster with words like stop, delay, arrest and restrict, which seems to correspond to the figurative sense of freeze. thus factoring out different senses allows mpro bert to give a more accurate and reasonable categorization. 7. discussion. one might argue that the superior performance of mpro bert embeddings is due simply to the increased total number of embeddings for this setting, compared to the static embeddings; because mpro bert has multiple representations per word form, it has more ‘chances’ to correctly categorize each word. we do find that the mpro bert clusters have more members on average: the average ground truth class has 55.5 members, while for our optimal mpro bert results, the average cluster size was 96.0 words. on average, one word form appeared in 3.02 clusters, but only in 1.14 ground truth classes. all else being equal, the larger cluster size and the fact that words appear in more clusters should lead to higher recall scores and lower precision. in fact, we find that mpro bert embeddings have higher precision and higher recall scores than the static embeddings, confirming that the difference we see between mpro bert and static embeddings is not merely a fluke; sense-level embeddings really do seem to better capture broad semantic categories better. in fact, mpro bert might be better at modeling human participants than suggested by our f1 measure. this is because mpro bert tends to capture more distinct senses per word than human participants did, as human participants generally focused on a single sense when categorizing. for example, the word form jump occurs in one mpro bert cluster corresponding to violence (jump#attack), another cluster corresponding to physical movement (jump#hop), and a third one related to change (jump#increase). in the ground truth data, jump only occurs once, in a class related to physical movement. perhaps this is the most salient sense of the word jump, and therefore participants were more likely to be thinking of this sense during the word sorting task and ignore its other possible senses. our f1 score is therefore a conservative measure of the success of mpro bert in modeling human categorization exclusively from linguistic input, in that it counts these other non-dominant senses of jump against mpro bert in our evaluation. however, the fact that embeddings for jump were assigned three separate clusters is not necessarily a weakness: the mpro bert clusters are more thorough as they represent each sense of the word separately and appropriately assign them to separate clusters. mpro bert clusters correspond to all possible ways of classifying sets of contexts (sense) words appear in to derive semantic categories; human participants seem to be more likely to grab onto fewer or even only one of set of contexts (sense). as the previous example suggests, our conservative f1 scores may not give a full picture of the quality or reasonableness of the word embedding clusters, as it abstracts away from the fact that human participants might only attend to a few or even one sense of words they semantically categorize even when those words are presented in isolation. our current results also abstract away from an important aspect of categorization, namely its flexibility and context-dependency: proceedings of elm 2: 278-287, 2023 elizabeth soper and jean-pierre koenig: modeling the role of polysemy in verb categorization. 284 https://doi.org/10.3765/elm https://www.elm-conference.net/ categorization is a relatively flexible task. there may be many possible criteria for sorting a group of words, especially when given such a large set of words to sort (tversky 1977, barsalou 1982). the flexibility of human categorization might explain the low inter-annotator agreement on the task; majewska et al. (2021) measured the agreement between two initial test participants on the verb sorting task, which was just 0.400 b-cubed score. this low agreement suggests that humans don’t perform very consistently in creating broad semantic categories from a large group of words. even within a single participant, categories were not created based on consistent criteria. as a result, it’s possible for induced categories from word embeddings to be reasonable, but still correlate poorly with our ground truth data. to fully determine how much semantic knowledge can be extracted from linguistic input, it is critical for models based on word embeddings to be able to mimic the flexibility and contextual-dependence of human categorization. 8. conclusion and future work. majewska et al. (2021) found that static word2vec embeddings were better at predicting human categorization of verbs than contextual bert embeddings. in this paper, we challenged this result, comparing sense-specific embeddings against previously evaluated static representations, and found that the rich, sense-specific information present in bert allows it to excel at predicting semantic categories. on a more general level, these results suggest that linguistic input encodes a great deal of information about semantic categories, independently of other perceptual input that humans receive, and that this category information can be extracted from embedding models which are trained on linguistic data alone. these results also suggest that word meanings that derive from linguistic input are better able to model human categorization if they correspond to semantically similar sets of contexts (the distributional equivalent of word senses). more broadly, our research supports the words-as-cues theory put forth in elman (2009) in that it is similar sets of contexts of use that best predict human categorization. future work is needed, though, to extend this analysis to nouns, which may behave differently, as well as to further explore the role of language in forming semantic categories, and whether models like those discussed here can model the flexible, goal-dependent nature of human categorization. references aharoni, roee & yoav goldberg. 2020. unsupervised domain clusters in pretrained language models. in proceedings of the 58th annual meeting of the association for computational linguistics, 7747–7763. https://doi.org/10.18653/v1/2020.acl-main.692. barsalou, lawrence w. 1982. context-independent and context-dependent information in concepts. memory & cognition 10(1). 82–93. https://doi.org/10.3758/bf03197629. bnc consortium. 2007. british national corpus. oxford text archive core collection . bommasani, rishi, kelly davis & claire cardie. 2020. interpreting pretrained contextualized representations via reductions to static embeddings. proceedings of the 58th annual meeting of the association for computational linguistics 4758–4781. https://doi.org/10.18653/v1/2020.acl-main.431. brocher, andreas, jean-pierre koenig, gail mauner & stephani foraker. 2018. about sharing and commitment: the retrieval of biased and balanced irregular polysemes. language, cognition, and neuroscience 33(4). 443–466. https://doi.org/10.1080/23273798.2017.1381748. chronis, gabriella & katrin erk. 2020. when is a bishop not like a rook? when it’s like proceedings of elm 2: 278-287, 2023 elizabeth soper and jean-pierre koenig: modeling the role of polysemy in verb categorization. 285 https://doi.org/10.3765/elm https://www.elm-conference.net/ a rabbi! multi-prototype bert embeddings for estimating semantic relationships. proceedings of the 24th conference on computational natural language learning 227–244. https://doi.org/10.18653/v1/2020.conll-1.17. collins, allan m. & elizabeth f. loftus. 1975. a spreading-activation theory of semantic processing. psychological review 82(6). 407–428. https://doi.org/10.1037/0033-295x.82.6.407. devlin, jacob, ming-wei chang, kenton lee & kristina toutanova. 2019. bert: pre-training of deep bidirectional transformers for language understanding. proceedings of the 2019 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) 4171–4186. https://doi.org/10.18653/v1/n19-1423. elman, jeffrey l. 2009. on the meaning of words and dinosaur bones: lexical knowledge without a lexicon. cognitive science 33(4). 547–582. https://doi.org/10.1111/j.15516709.2009.01023.x. fares, murhaf, andrey kutuzov, stephan oepen & erik velldal. 2017. word vectors, reuse, and replicability: towards a community repository of large-text resources. proceedings of the 21st nordic conference on computational linguistics, nodalida, 22-24 may 2017, gothenburg, sweden 131. 271–276. giulianelli, mario, marco del tredici & raquel fernández. 2020. analysing lexical semantic change with contextualised word representations. proceedings of the 58th annual meeting of the association for computational linguistics 3960–3973. https://doi.org/10.18653/v1/2020.acl-main.365. landauer, thomas k & susan t dumais. 1997. a solution to plato’s problem: the latent semantic analysis theory of acquisition, induction, and representation of knowledge. psychological review 104(2). 211. https://doi.org/10.1037/0033-295x.104.2.211. liu, qianchu, diana mccarthy, ivan vulić & anna korhonen. 2019. investigating cross-lingual alignment methods for contextualized embeddings with token-level evaluation. proceedings of the 23rd conference on computational natural language learning (conll) 33–43. https://doi.org/10.18653/v1/k19-1004. majewska, olga, diana mccarthy, jasper jf van den bosch, nikolaus kriegeskorte, ivan vulić & anna korhonen. 2021. semantic data set construction from human clustering and spatial arrangement. computational linguistics 47(1). 69–116. https://doi.org/10.1162/coli a 00396. marvel, aron & jean-pierre koenig. 2015. event categorization beyond verb senses. proceedings of the 11th workshop on multiword expressions (mwe 2015) https://doi.org/10.3115/v1/w150913. mikolov, tomas, kai chen, greg corrado & jeffrey dean. 2013. efficient estimation of word representations in vector space. 1st international conference on learning representations, iclr 2013 workshop track proceedings 1–12. miller, george a. 1995. wordnet: a lexical database for english. communications of the acm 38(11). 39–41. https://doi.org/10.1145/219717.219748. parker, robert, david graff, junbo kong, ke chen & kazuaki maeda. 2011. english gigaword fifth edition. linguistic data consortium. pennington, jeffrey, richard socher & christopher d manning. 2014. glove: global vectors for proceedings of elm 2: 278-287, 2023 elizabeth soper and jean-pierre koenig: modeling the role of polysemy in verb categorization. 286 https://doi.org/10.3765/elm https://www.elm-conference.net/ word representation. proceedings of the 2014 conference on empirical methods in natural language processing (emnlp) 1532–1543. https://doi.org/10.3115/v1/d14-1162. pereira, francisco, samuel gershman, samuel ritter & matthew botvinick. 2016. a comparative evaluation of off-the-shelf distributed semantic representations for modelling behavioural data. cognitive neuropsychology 33(3-4). 175–190. https://doi.org/10.1080/02643294.2016.1176907. peters, matthew e., mark neumann, mohit iyyer, matt gardner, christopher clark, kenton lee & luke zettlemoyer. 2018. deep contextualized word representations. proceedings of the 2018 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 1 (long papers) 2227–2237. https://doi.org/10.18653/v1/n18-1202. pustejovsky, james. 1998. the semantics of lexical underspecification. folia lingüı́stica: acta societatis linguisticae europaeae 32(3). 323–348. https://doi.org/10.1515/flin.1998.32.3-4.323. sia, suzanna, ayush dalmia & sabrina j mielke. 2020. tired of topic models? clusters of pretrained word embeddings make for fast and good topics too! proceedings of the 2020 conference on empirical methods in natural language processing (emnlp) 1728–1736. https://doi.org/10.18653/v1/2020.emnlp-main.135. soler, aina gari & marianna apidianaki. 2021. let’s play mono-poly: bert can reveal words’ polysemy level and partitionability into senses. transactions of the association for computational linguistics 9. 825–844. https://doi.org/10.1162/tacl a 00400. tversky, amos. 1977. features of similarity. psychological review 84(4). 327–352. https://doi.org/10.1037/0033-295x.84.4.327. vulić, ivan, simon baker, edoardo maria ponti, ulla petti, ira leviant, kelly wing, olga majewska, eden bar, matt malone, thierry poibeau et al. 2020. multi-simlex: a large-scale evaluation of multilingual and crosslingual lexical semantic similarity. computational linguistics 46(4). 847–897. https://doi.org/10.1162/coli a 00391. wolf, thomas, lysandre debut, victor sanh, julien chaumond, clement delangue, anthony moi, pierric cistac, tim rault, rémi louf, morgan funtowicz et al. 2020. transformers: state-of-the-art natural language processing. proceedings of the 2020 conference on empirical methods in natural language processing (emnlp): system demonstrations 38–45. https://doi.org/10.18653/v1/2020.emnlp-demos.6. proceedings of elm 2: 278-287, 2023 elizabeth soper and jean-pierre koenig: modeling the role of polysemy in verb categorization. 287 https://doi.org/10.3765/elm https://www.elm-conference.net/ the rise and particularly fall of presuppositions: evidence from duality in universals remus gergel, maike puhl, simon dampfhofer & edgar onea* abstract. at the center of this paper is the question whether presuppositions are more likely to be gained or lost in the process of language change. we offer a new experimental method that aims at ascertaining the re-learning speed of potentially presuppositional items based on nonce words and which integrates certain factors of change such as social prestige in an artificial but clearly contextualized set-up. the meaning targeted is of a quantifier meaning ‘both’ with speakers of german and the initial results point to higher ease of losing rather than incorporating the presupposition, but with an interesting resilience after a critical questioning of presuppositional status. keywords. diachronic semantics; presuppositional quantifiers; reinterpretation times; semantic language processing 1. introduction. in the theoretical literature on language change and historical linguistics, one fundamental question pertains to the ways in which different layers of meaning interact and are transformed, or even vanish. the classical view is, for instance, that implicatures conventionalize and thus become lexical entries, while disappearing as implicatures (see, e.g., traugott & dasher 2002, eckardt 2006 for discussions). other central aspects of meaning at the semantics-pragmatics interface have, however, been less studied. crucially for the current enterprise, the question how lexical items come to be presupposition triggers or lose this property has been much less discussed on a principled basis, although case studies on presuppositional items certainly exist (see beck, berezovskaya & pflugfelder 2009, schwenter & waltereit 2010, los & komen 2012, beck & gergel 2015, gergel, blümel & kopf 2016, gergel & kopf-giammanco 2021, carlier & lamiroy 2018, among others). essentially, there are two types of approaches to our central question that have been discussed, even if they have been developed on a very different set of examples. on the one hand, eckardt (2009) proposes a principle of avoid pragmatic overload which incorporates presuppositions and essentially states that in certain critical situations in which there are too many side-messages (including presuppositions and implicatures) speakers will drop some. this can be understood as presuppositional loss. gergel (2020), on the other hand, focuses on the increased presence of the triggers of presuppositions across different contexts over time and suggests a diachronic corollary of the synchronic principle of maximize presuppositions, which comes down to an increase as far as the signaling of presuppositional meaning components goes. however, even if historical linguistics can elucidate the historical path of development for any individual presupposition trigger, the question how and why, in terms of cognitive underpinning, such shifts come about still remains an open and important one. any theory put forth by historical linguistics on this matter would come with particular cognitive predictions that would need to be borne out generally, irrespective of the peculiarities of the historical processes. thus, one would *we thank the audience of elm2 for their feedback. we also thank alexandre cremers for helpful methodological discussion regarding the statistical analysis of our results at a workshop in vilnius financed by the arqus network. authors: remus gergel, saarland university (remus.gergel@uni-saarland.de), maike puhl, saarland university (maike.puhl@uni-saarland.de), simon dampfhofer, university of graz (simon.dampfhofer@edu.uni-graz.at) & edgar onea, university of graz (edgar.onea-gaspar@uni-graz.at). proceedings of elm 2: 72-82, 2023 c©2023 remus gergel, maike puhl, simon dampfhofer and edgar onea published by the lsa with permission of the author(s) under a cc by license. 72 https://doi.org/10.3765/elm https://www.elm-conference.net/ expect to be able to reproduce and test mechanisms of language change in the lab to at least some extent. such experimental systematic research is non-existent to our knowledge for the classical areas of presupposition triggers; however, some studies that can be considered comparable in the broader sense are available. this line of research falls largely, of course, into the classical labovian idea of experimentally exploiting present reactions of speakers to uncover processes that can be relevant for language change more generally. and just as naturally, our approach shares more with recent attempts to explain paths of change in the area of meaning (rather than sound change or morphosyntax as in labovian studies), such as zhang, piñango & deo (2018), fedzechkina & roberts (2020), fuchs, deo & piñango (2020), gergel, kopf-giammanco & puhl (2021), puhl & gergel (2022). returning to the general question regarding presuppositions, our paper addresses it by considering one single case of a presupposition trigger, the lexical item corresponding to the meaning of the quantifier both. on the classical view (see heim & kratzer 1998 for discussion), this quantifier is similar to a universal quantifier such as all, but it has the presuppositional restriction that it can only be used appropriately in cases in which the restrictor set has the cardinality two. the question we ask is this: how easy is it for speakers to start with the presuppositional meaning of both and learn (and adapt to) a new meaning in an experimental scenario that emulates language change in certain ways, corresponding to the meaning of all, as compared to starting with the non-presuppositional item all and learning and adapting to a new presuppositional meaning in the experimental setup corresponding to the meaning of both? thus, this would give us a direct way to compare the ease at which an item can pick up or lose a presupposition in a language change situation. in what follows, we describe our key experiment (section 2) together with a relevant replication (section 3) with their respective methods and results, before returning to a more general discussion in section 4. 2. experiment 1. experiment 1 was designed to investigate whether it would be easier to reinterpret both as all (notation: both →all) or vice versa (all→both). to avoid any influence of previous knowledge of the actual words both or all (i.e., their german variants), participants were taught a nonce word, gure, instead, which would mean either both or all. during the experiment, participants were asked to imagine visiting a fictitious community (german speaking diaspora in the us). they were exposed to two types (roughly correlating with generations) of native speakers. older speakers would use gure in its original meaning, while younger and up-to-date speakers would use the reinterpreted meaning (opposite quantifier; e.g. both instead of the original all). 2.1 method. we recruited 25 native speakers of german by advertising on university newsletters and on university related social media groups (11m/14f; mean age 23.1, sd 3.2). they were financially remunerated for their participation. the experiment was conducted on-site at the university of graz. participants were separated into two groups; one group would learn that gure originally had the meaning of both (13 participants), the other learnt that gure originally had the meaning of all (12 participants). in addition to gure, participants were taught two filler noncewords whose meaning had no presuppositional component. the experiment was split into a training phase and a test phase. during the training phase, participants were told that they should imagine being accompanied by an older native speaker as well as a young non-native friend. they were then shown images on a computer screen and heard sentences containing a nonce word, produced by the non-native person describing the situation. after this, the old person would tell participants whether the sentence was true in the situation presented. if the sentence was not true, the old person would, in addition, provide a reason why it proceedings of elm 2: 72-82, 2023 remus gergel, maike puhl, simon dampfhofer and edgar onea: the rise and particularly fall of presuppositions. 73 https://doi.org/10.3765/elm https://www.elm-conference.net/ was false. the older speaker didn’t adhere to a maximize presupposition principle; in situations where two out of two items were relevant, they accepted both the meaning of both and all as true. spoken stimuli of native speakers were produced in a version of the saarland dialects (a remote variant of mosel-franconian), which would sound exotic in the south-eastern austrian region of graz where the study was conducted. the non-native friend was a speaker from graz who spoke a rather neutral dialect similar to standard german. a sample item is given in figure 1, with an added english translation. there the contribution of the old speaker differed in the all→both and both→all conditions respectively. alter sprecher: “sag mal, sarah, was siehst du in dem korb? junge sprecherin: “gure äpfel sind rot.” alter sprecher: “sehr gut.” (all→both) “nein, das stimmt nicht. es sind mehr als zwei äpfel.” (both→all) old speaker: “tell me, sarah, what do you see in this basket?” young speaker: “gure apples are red.” old speaker: “very good.” (all→both) “no, that’s not right. there are more than two apples.” (both→all) figure 1: sample item, training phase after three training items each, participants were asked to rate the truth of five sentences on a binary scale. this was done in order to verify the success of the training. in addition to judgments, we also measured the reaction times of participants. participants received written feedback after each judgment whether their choice was correct, as well as an explanation of their mistake in case they were wrong. afterwards, they were exposed to additional blocks consisting of training items and test judgments. in total, the training phase consisted of 5 gure blocks and 2 filler blocks. all training items were presented in a fixed order. in the second phase of the experiment (test phase), participants were asked to imagine visiting a reunion of younger members of the community. the young native speakers would use gure in its reinterpreted meaning, i.e., both →all or all→both. participants were introduced to a friend f who had been abroad for some time and was thus not up to date with current language developments, and a high prestige competent local speaker s. they were then shown images and heard sentences. participants were told that they didn’t know which of the people attending the reunion had produced the sentences (which means that the speaker could either have high or low competence regarding the reinterpreted meaning). after hearing the sentences, participants were asked to rate their agreement for the sentence on a scale from 1 to 10. after each item, participants proceedings of elm 2: 72-82, 2023 remus gergel, maike puhl, simon dampfhofer and edgar onea: the rise and particularly fall of presuppositions. 74 https://doi.org/10.3765/elm https://www.elm-conference.net/ then read a short dialogue between f and s commenting on the situation. both would mention if they thought that the speaker had made a mistake but would not explain why they thought the sentence was wrong. participants were shown 18 items containing gure and 6 filler items which were separated into five blocks and randomized within the blocks. the meaning of some fillers had changed compared to the meaning used by the older speaker during training, while the meaning of others had not. a sample item is given in figure 2. sprecher: “gure roten äpfel sind faul.” both→all: f: “hast du den dritten faulen apfel gesehen?” s: “wieso sagst du das, lara? das stimmt doch, er hat ‘gure’ gesagt.” all→both: f: “mist. wie viele grüne haben wir?” s: “das passt aber nicht. "gure" klingt nach etwas, was hier höchstens meine oma sagen würde!” speaker: “gure red apples are rotten.” both→all: f: “did you see the third rotten apple?” s: “why do you say that, lara? he said ‘gure’, he was right.” all→both: f: “damn. how many green ones do we have?” s: “that’s not right. ‘gure’ sounds like something my grandma would say here!” figure 2: sample item, experiment phase 2.2. results. all participants completed the training successfully, meaning that they answered with the expected judgments in the last blocks. also, judgments of items containing those fillers whose meaning had not changed compared to the training phase were as expected, i.e., they did not change significantly compared to judgments during training. we analyzed reaction times and judgments of the test phase by fitting data for all relevant items (i.e., items where either both would be correct but all would not be, or vice versa) with linear mixed models in r (r core team 2021) using the package lme4 (bates, mächler, bolker & walker 2015).2 in particular, we used the order 2 there is a methodological question whether our 10-point likert scales fulfill the requirements for applying the lmer function from the lme4 package. after all, one could argue that our data could in principle exhibit ceiling effects and thus should rather be analyzed with the cumulative link mixed model such as the clmm function from the ordinal package (christensen 2019). we generally believe that a 10-point likert scale, on which the extremes were rarely reached at all in the experiment, does qualify as discrete rather than ordinal data much more than in the case of a 4or 7-point likert scale. thus, treating our results as involving ordinal data would lead to loss of information. we did perform cumulative link models with the same parameters in all cases and found similar trends, but generally less proceedings of elm 2: 72-82, 2023 remus gergel, maike puhl, simon dampfhofer and edgar onea: the rise and particularly fall of presuppositions. 75 https://doi.org/10.3765/elm https://www.elm-conference.net/ of presentation of the items, the group variable (reflecting the both→all vs. all→both) conditions, and random slopes and intercepts for items for both reaction times and judgments, as reflected by the formulae in (1) and (2) below: (1) reactiontime ~ order * group + (1 + group | item) + (1 + order | item) (2) judgment ~ order * group + (1 + group | item) + (1 + order | item) results are shown in figure 3 and table 1 for reaction times and in figure 4 and table 2 for judgements (using r packages lüdecke 2021 for plotting and fox, price & weisberg 2019 for anova). as can be seen, we found a significant main effect for order in both cases, and an interaction between order and group for reaction times, and a main effect for groups in the case of judgments. figure 3: results reaction time, experiment 1 chisq df pr(>chisq) order 14.4387 1 0.0001448 *** group 3.3893 1 0.0656210 . order:group 13.0456 1 0.0003037 *** signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 table 1: anova for reaction time model, experiment 1 in the plots, “reactiontime” indicates the time between the end of the audio stimulus (i.e. heard sentence) and the participants reaction (button press). “judgment” represents acceptance of a sentence given a situation on a scale from 1 to 10. for better comparability, we inverted the scale of judgments for the all→both-group so that it goes from 10 to 1 instead of 1 to 10. this way, low judgment values for both groups indicate judgments according to the original meaning of gure learnt during training, while high values indicate acceptance of the reinterpreted meaning. on the significant effects. in the case of experiment 1, the cumulative link mixed model shows a trend of order (p = 0.051), no significant effect of group, and no significant interaction; the dataset is too small to support both main effects in addition to interaction. in contrast, a cumulative link mixed model taking into account order and group but assuming no interaction does show significant effects of order (p = 0.00608) as well as group (p = 0.00233). while we maintain that we deem linear mixed models to be a more appropriate method of analysis for our data, we acknowledge this to be a debatable issue. proceedings of elm 2: 72-82, 2023 remus gergel, maike puhl, simon dampfhofer and edgar onea: the rise and particularly fall of presuppositions. 76 https://doi.org/10.3765/elm https://www.elm-conference.net/ x-axis, order represents participants’ progress during the experiment phase (order 0 is the start of the experiment, and order 1 the end). figure 4: results judgments, experiment 1 chisq df pr(>chisq) order 7.0845 1 0.007775 ** group 6.6758 1 0.009773 ** order:group 1.6200 1 0.203098 signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 table 2: anova for judgment model, experiment 1 2.3 discussion. looking at the judgments data in more detail, we can observe that some participants would accept the reinterpreted meaning rather quickly while others wouldn’t accept the changed meaning at all. the result is a split between participants with some responding with high values of judgment towards the end of the experiment and some with low values. this trend was not group specific; both groups showed a similar split. it is likely due to a general reluctance of some people to accept deviations from a meaning they had been explicitly taught before. more importantly, we take the combination of the analysis of reaction times and judgments to have an added value regarding the interpretation of the results. while reaction times only show an interaction suggesting that participants got quicker over time in the both→all group more than in the other group, and the judgments only show that both→all was also judged better overall (with only a marginal tendency of an interaction), we believe that these findings provide strong evidence for a quicker and better adaptation to language change in the both→all direction. 3. experiment 2. to get more reliable information on whether participants actually understood the meaning of gure as a presupposition (as opposed to truth-conditional meaning), we conducted follow-up experiments which replicated the method of experiment 1 but added presupposition tests (family of sentences tests). in addition, replicating experiment 1 appeared to be a useful step overall, given that our findings in the first experiment can be widely considered explorative and thus less reliable despite statistical significance. 3.1 method. we recruited 24 native speakers of german, again by advertising on university newsletters and on university related facebook groups (16f/8m; mean age 24.4, sd 4.2). they were financially remunerated for their participation. due to a technical error, two participants’ data proceedings of elm 2: 72-82, 2023 remus gergel, maike puhl, simon dampfhofer and edgar onea: the rise and particularly fall of presuppositions. 77 https://doi.org/10.3765/elm https://www.elm-conference.net/ were unusable, resulting in 22 usable datasets (16f/6m; mean age 23.8, sd 3.3), 11 per group. the experiment was conducted on-site at the university of graz. the setup of the training and test phases was identical to experiment 1. in addition, we added presupposition test surveys both after the training phase and after the test phase. in each survey, participants were shown 15 items consisting of a sentence from the family of sentences paradigm as well as a question about that sentence. five out of those questions tested the presupposition of there being exactly two relevant items, while the other 10 questions tested other conditions. participants were then asked to answer the question on a binary scale (“yes” or “no”). assuming s to be any sentence, we used five different types of embedding: if-s, not-s, maybe-s, thinks-thats, and says-that-s. an example for an if-s embedding together with a translation is shown in figure 5. jessica sagt: „wenn gure milchpackungen leer sind, müssen wir neue kaufen.“ kannst du dann davon ausgehen, dass jessica der meinung ist, dass es genau zwei milchpackungen gibt? jessica says: “if gure cartons of milk are empty, we have to buy new ones.” can you then assume that jessica thinks that there are exactly two cartons of milk? figure 5: example item presupposition test, experiment 2 3.2 results. again, all participants completed the training successfully. the statistical analysis we used for the two experiments were identical. figure 6: reaction time results, experiment 2 chisq df pr(>chisq) order 11.6035 1 0.0006583 *** group 15.2507 1 9.414e-05 *** order:group 2.4792 1 0.1153595 signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 table 3: anova for reaction time model, experiment 2 judgment and reaction time results roughly mirror the results from experiment 1, which was expected due to the analogous method; these are shown in figure 6 and figure 7 graphically and proceedings of elm 2: 72-82, 2023 remus gergel, maike puhl, simon dampfhofer and edgar onea: the rise and particularly fall of presuppositions. 78 https://doi.org/10.3765/elm https://www.elm-conference.net/ the significance tests for the linear models in table 3 and table 4, respectively. this time, however, for both reaction times and judgments, we observed significant main effects of group and order but no significant interaction.3 figure 7: judgment results, experiment 2 chisq df pr(>chisq) order 20.3314 1 6.512e-06 *** group 18.4323 1 1.760e-05 *** order:group 1.3985 1 0.237 signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 table 4: anova for judgment model, experiment 2 we analyzed the presupposition tests by comparing the mean results of both surveys for each group. the results are shown in figure 8, which represents the numerical values also summarized in table 5. figure 8: presupposition test acceptance, experiment 2 as expected, there is a considerable difference between the two groups after the training phase, since only the both-group actually acquired a presupposition, whereas the results strongly overlap after the test phase. even if visually obvious, we checked for statistical significance using a 3 unlike experiment 1, here the cumulative link mixed model taking into account both main effects as well as interaction showed significant effects of order (p = 0.00886) and group (p = 0.03826), and no significant interaction. proceedings of elm 2: 72-82, 2023 remus gergel, maike puhl, simon dampfhofer and edgar onea: the rise and particularly fall of presuppositions. 79 https://doi.org/10.3765/elm https://www.elm-conference.net/ generalized linear mixed model with group and phase as predictors and random intercepts for items. this revealed a highly significant interaction, as witnessed in table 6. after training after test all→both 0.02±0.13 0.45±0.50 both→all 0.70±0.46 0.44±0.50 table 5: numerical results for presupposition tests chisq df pr(>chisq) phase 0.0040 1 0.9492653 group 5.3978 1 0.0201618 * phase:group 15.1168 1 0.0001011 *** signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 table 6: anova for presupposition test model, experiment 2 3.3 discussion. experiment 2 confirmed the findings of experiment 1 and additionally showed that participants acquired the presuppositional meaning of gure in the training phase and lost it during the test phase for the both→all group, whereas in the other condition they started without a presuppositional meaning for gure after the training phase and acquired it to some degree during test phase. the high standard deviation for the presupposition test after the test phase can be explained by the observed split in participants’ acceptance of the reinterpreted meaning. this would result in mean judgments near the middle of the total range with high standard deviations, though there were clear groups of participants who achieved a very high level of learning in both groups. 4. overall discussion. our results indicate that it is easier to perform reinterpretation in a contextualized language change situation in such a way as to lose rather than gain a presupposition. this is witnessed by a quicker and better learning process in terms of adaptation to ongoing language change in the case of presuppositional loss as compared to acquiring a presuppositional item. this has been ascertained in experiment 1 in an exploratory manner and confirmed in experiment 2. in addition, experiment 2 also confirmed that the acceptabilities measured in both experiments are due to presuppositional effects proper and not some alternative prototypicality effects. however, our results are limited to one single presuppositional item. clearly, such preliminary results must be refined and validated further, especially using further types of presupposition triggers. from an experimental perspective, the paradigm we have set up invites both extensions and clarifications. while we sought contextualization, including with respect to the simulation of sociolinguistic factors that are relevant for change, it should be obvious that an artificial set-up can always be improved in multiple ways. a further factor that is inherent to most presuppositional conditions will be that they impose additional restrictions (e.g., the duality restriction in our case). this opens up the possibility that a reinterpretation leading towards their loss is also one that leads to a generalization. while in actual change both generalization and specialization can occur, this is experimentally a path that we plan to control in future experiments, also independently of purely presuppositional items. the fact that we only studied one item equally invites further research. from a language change perspective, notice that we did not start from the expectation of replicating a particular change, but rather from the idea of distilling relevant cognitive behavior proceedings of elm 2: 72-82, 2023 remus gergel, maike puhl, simon dampfhofer and edgar onea: the rise and particularly fall of presuppositions. 80 https://doi.org/10.3765/elm https://www.elm-conference.net/ patterns that can provide some of the fundamental building blocks in the construction of change.4 the complexities of actualized changes can require experimental checking with respect to multiple dimensions. places in which the potential weakening of the duality restriction we have studied here may show are, as discussed in gergel (2022), developments of the quantified both in early middle english times in which it was reinforced by the numeral two as well as similar developments in the histories of some romance languages in which versions of the original latin quantifier ambo have equally been reinforced by numerals as e.g., in developments of romanian amândoi or spanish ambos dos (while the latter is prescriptively often argued against, the former is unmarked). more generally, however, notice that with respect to theories of actual change our experimental results came closest to eckardt’s assumptions about the possibility of losing presuppositions. but the interesting suggestions from eckardt’s work, too, can and must be qualified from the present perspective, as we did not introduce any particularly critical and burdensome situations in which speakers would be overwhelmed with multiple inferences. therefore, as it stands, our current result detaches potential presuppositional loss from a principle such as avoid pragmatic overload. 5. references. bates, douglas, martin mächler, ben bolker & steven walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67,1. 1–48. https:// beck, sigrid & remus gergel. 2015. the diachronic semantics of english again. natural language semantics 23,3. 157–203. beck, sigrid, polina berezovskaya, & katja pflugfelder. 2009. the use of again in 19th‐century english versus present‐day english. syntax, 12,3. 193-214. carlier, anne & béatrice lamiroy. 2018. the emergence of the grammatical paradigm of nominal determiners in french and in romance: comparative and diachronic perspectives. canadian journal of linguistics/revue canadienne de linguistique, 63. 141-166. christensen, rune haubo bojesen. 2019. ordinal: regression models for ordinal data. https://cran.r-project.org/web/packages/ordinal/. eckardt, regine. 2006. meaning change in grammaticalization. oxford: oup. eckardt, regine. 2009. apo: avoid pragmatic overload. in jacqueline visconti and maj-britt mosegaard hansen (eds.), current trends in diachronic semantics and pragmatics, 21–42. london: emerald. fedzechkina, masha & gareth roberts. 2020. learners sacrifice robust communication as a result of a social bias. proceedings of the 42nd annual meeting of the cognitive science society, 2281–2287. fox, john, brad price, & sanford weisberg. 2019. car: companion to applied regression. https://cran.r-project.org/web/packages/car/. fuchs, martín, ashwini deo & maría mercedes piñango. 2020. the progressive-to-imperfective shift. contextually determined variation in rioplatense, iberian, and mexican altiplano spanish. in alfonso morales-front, michael j. ferreira, ronald p. leow and cristina sanz (eds.), hispanic linguistics: current issues and new directions, 119–136. amsterdam: benjamins. 4 for so-called atoms from the perspective of semantic theory, see e.g., grosz (2022) for a recent take and the literature cited there. our focus here is of course on cognitively verified building blocks. proceedings of elm 2: 72-82, 2023 remus gergel, maike puhl, simon dampfhofer and edgar onea: the rise and particularly fall of presuppositions. 81 https://doi.org/10.3765/elm https://www.elm-conference.net/ gergel, remus. 2020. sich ausgehen: actuality entailments and further notes from the perspective of an austrian german motion verb construction. proceedings of the linguistic society of america, 5(2). 5–15. https://doi.org/10.3765/plsa.v5i2.4790. gergel, remus. 2022. cyclicity effects in the development of presuppositions. ms. saarland university. gergel, remus, andreas blümel & martin kopf. 2016. another heavy road of decompositionality: notes from a dying adverb. university of pennsylvania working papers in linguistics 22, 1. gergel, remus & martin kopf-giammanco. 2021. ‘sich ausgehen’: on modalizing go constructions in austrian german. canadian journal of linguistics/revue canadienne de linguistique 66,2. 141-190. gergel, remus, martin kopf-giammanco & maike puhl. 2021. simulating semantic change: a methodological note. in proceedings of experiments in linguistic meaning (elm) 1, 184– 196. university of pennsylvania: lsa publications. https://doi.org/10.3765/elm.1.4869. grosz, patrick. 2022. scalarity as a meaning atom in wohl-type particles. in remus gergel, ingo reich, augustin speyer (eds.), particles in german, english, and beyond, 243–268. amsterdam: benjamins. heim, irene & angelika kratzer. 1998. semantics in generative grammar. oxford: blackwell. los, bettelou & erwin komen 2012. clefts as resolution strategies after the loss of a multifunctional first position. in terttu nevalainen & elizabeth closs traugott (eds.), the oxford handbook of the history of english, 884–898. oxford: oup. lüdecke, daniel. 2021. sjplot: data visualization for statistics in social science. https://cran.r-project.org/package=sjplot. puhl, maike & remus gergel. 2022. final though. in remus gergel, ingo reich, augustin speyer (eds.), particles in german, english, and beyond, 177–208. amsterdam: benjamins. r core team. 2021. r: a language and environment for statistical computing. vienna, austria. https://www.r-project.org/. schwenter, scott & richard waltereit. 2010. presupposition accommodation and language change. in kristin davidse, lieven vandelanotte, and hubert cuyckens (eds.), subjectification, intersubjectification and grammaticalization, 75–102. berlin: mouton de gruyter. traugott, elizabeth closs & richard b. dasher. 2012. regularity in semantic change cambridge: cambridge university press. zhang, muye, maria mercedes piñango & ashwini deo. 2018. real-time roots of meaning change: electrophysiology reveals the contextual-modulation processing basis of synchronic variation in the location possession domain. 40th annual conference of the cognitive science society, 2783–2788. proceedings of elm 2: 72-82, 2023 remus gergel, maike puhl, simon dampfhofer and edgar onea: the rise and particularly fall of presuppositions. 82 https://doi.org/10.3765/elm https://www.elm-conference.net/ crosslinguistic differences on the present perfect puzzle: an experimental approach martín fuchs & martijn van der klis* abstract. in this paper, we analyze how different temporal and referential properties of past-referring adverbials –specifically, hodiernality and deixis– are partially responsible for the crosslinguistic distribution of past and perfect markers across dutch, spanish, and english. to that end, we conducted an acceptability judgment task, where 160 subjects per language rated context-sentence pairs that display either a past or a perfect marker, and a temporal adverbial that is: (i) either temporally close to or temporally far from the speech time, and (ii), either deictic or not deictic. results show that: (a) dutch allows for its perfect marker to combine with any past-referring temporal adverbial, (b) spanish only allows its perfect marker to combine with adverbials that locate the event temporally close to speech time, regardless of deixis, and (iii) that english prefers its past marker in all past-referring situations, but allows its perfect to combine with adverbials that are both deictic and temporally close to speech time, particularly when the adverb specifies an interval that is included in the day of utterance (e.g., this morning), as opposed to adverbs that describe an interval that includes it (e.g., this month). keywords. perfect; past; grammatical aspect; temporal adverbials; deixis; hodiernality; acceptability judgments 1. introduction. the distribution of the present perfect1and the simple past in english is said to encompass a distinction between past events that have some current relevance or might continue into the present –which are expressed using the present perfect–, and past events that started and ended before utterance time –expressed with the simple past. this contrast is especially observable when the events at issue are anchored to a definite or specific time in the past. since klein (1992), it is well known that the present perfect marker cannot combine with a temporal adverbial referring to the past, so that the simple past has to be used instead, as in (1): (1) chris *has left / left today at three o’clock. (adapted from klein 1992: 546, (ex.45)) in the sentence in (1), the specification of the past time reference (‘today at three o’clock’) makes the use of the present perfect form has left impossible, so that the only available option is the simple past form left. this incompatibility has been referred to in the literature as the present perfect puzzle, since both the present perfect and the simple past are able to locate an event in * we would like to thank the organizers of elm 2 –for developing such a great venue for the exchange of ideas on linguistic meaning and experimentation– and its audience, for helpful feedback and comments. we would also like to thank the audiences at lsrl 52 and cls 58 for valuable suggestions about the research presented here. the time in translation (tint) team at utrecht university has been fundamental in the pursuit of this project, too, and we thank all of its members for their critical engagement with our work. finally, we thank our funding agency, nwo grant 36080-070, without whose support the work reported here would not have been possible. all errors and omissions remain our own. authors: martín fuchs, utrecht institute of linguistics, utrecht university (m.fuchs@uu.nl) & martijn van der klis, utrecht institute of linguistics, utrecht university (m.h.vanderklis@uu.nl). 1 we use italics to indicate language-specific forms, and small caps for crosslinguistic marker types comprising a set of language-specific forms (e.g., present perfect in english, passé composé in french, and pretérito perfecto compuesto in spanish, among others, correspond to the perfect crosslinguistically). we reserve plain text with initial capitalization to refer to meanings. proceedings of elm 2: 61-71, 2023 c©2023 martı́n fuchs and martijn van der klis published by the lsa with permission of the author(s) under a cc by license. 61 https://doi.org/10.3765/elm https://www.elm-conference.net/ the past, making it somewhat surprising that one of these markers –the present perfect– is ultimately incompatible with past-referring adverbials. most solutions for this pattern invoke a reichenbachian framework (reichenbach 1947), where an additional time point –the reference time– is considered. according to reichenbach (1947), both the simple past and the present perfect situate the time of the event (e) before the time of speech (s), but the difference between these tense markers resides in the position of reference time (r). in the case of the simple past, the reference time (r) coincides with the event time (e), producing a e,r < s configuration, where < indicates precedence in time. conversely, the present perfect makes the reference time (r) coincide with speech time (s), such that the configuration becomes e < r,s. therefore, since temporal adverbials are argued to modify or target the reference time (r), the present perfect becomes incompatible with past-referring temporal adverbials, because these adverbs cannot target or modify a reference time that coincides with speech time. other western european languages, however, such as italian (squartini and bertinetto 2000), french (vet 1980, 1992) or german (musan 2002, schaden 2009), do allow their corresponding perfect markers (i.e., passato prossimo, passé composé, perfekt) to combine with past-referring adverbials, showing that the aforementioned constraint does not hold crosslinguistically. in a translation from (1), (2) illustrates this for italian, while (3) does so for french, and (4) for german:2 (2) chris è partito oggi alle tre. chris be.prs.3sg leave.pst.ptcp today at.the three lit. ‘chris has left today at three o’clock.’ (3) chris est parti aujourd’hui à trois heures. chris be.prs.3sg leave.pst.ptcp today at three hours lit. ‘chris has left today at three o’clock.’ (4) chris ist heute um drei uhr abgefahren. . chris be.prs.3sg today at three hours leave.pst.ptcp lit. ‘chris has left today at three o’clock.’ the dutch perfect marker, the voltooid tegenwoordige tijd (vtt, henceforth), is not affected by this constraint either, as (5a) shows. rather, as (5b) exemplifies, the perfect is said to be preferred over the dutch past, the onvoltooid verleden tijd (ovt, henceforth), in such contexts (de swart 2007, van der klis et al. 2022): (5a) chris is vandaag om drie uur vetrokken. chris be.prs.3sg today at three hours leave.pst.ptcp lit. ‘chris has left today at three o’clock.’ (5b) #chris vertrok vandaag om drie uur. chris leave.pst.3sg today at three hours lit. ‘chris left today at three o’clock.’ the dialectal varieties of peninsular spanish spoken in madrid and its surroundings appear to reflect an intermediate point in its availability to combine the spanish perfect marker –the pretérito perfecto compuesto– with past-referring temporal adverbials. these dialects allow the perfect to 2 for interlinear glosses, we use the abbreviation system and formatting conventions of the leipzig glossing rules, which can be found at https://www.eva.mpg.de/lingua/pdf/glossing-rules.pdf. additionally, we use prefl as in ‘pseudoreflexive’ to refer to the morphosemantic value of se in spanish that can appear with some intransitive verbs such as ir ‘to go’, resulting in a different meaning, as in irse ‘to leave’. proceedings of elm 2: 61-71, 2023 martı́n fuchs and martijn van der klis: crosslinguistic differences on the present perfect puzzle. 62 https://doi.org/10.3765/elm https://www.elm-conference.net/ express a past event anchored to a specific time in the past, as long as the event has occurred in a time that is included in the day of utterance (harris 1982, schwenter 1994, gonzález et al. 2019, fuchs & gonzález 2022). this contrast is shown in (6a-b) and (7a-b): (6a) chris se ha ido hoy a las tres. chris prefl have.prs.3sg leave.pst.ptcp today at the three lit. ‘chris has left today at three (o’clock).’ (6b) #chris se fue hoy a las tres. chris prefl leave.pst.pfv.3sg today at the three lit. ‘chris left today at three (o’clock).’ (7a) *chris se ha ido ayer. chris prefl have.prs.3sg leave.pst.ptcp yesterday lit. ‘chris has left yesterday.’ (7b) chris se fue ayer. chris prefl leave.pst.pfv.3sg yesterday lit. ‘chris left yesterday.’ as (6a) indicates, spanish allows its perfect to combine with temporal adverbials (hoy a las tres ‘today at three o’clock’) that create the relation e,r ⊆ day(s); that is, a relation in which the reference time (r) coincides with the event time (e), and both of these times are properly included within the day of the speech time (s). moreover, in these cases, spanish not only seems to allow the use of this marker, but also to prefer it over the spanish perfective past, the pretérito indefinido, as in (6b). conversely, when the event time (e) is anchored to a past reference time (r) before the day of utterance (s), creating a relation e,r < day(s), as in (7), with the adverb ayer ‘yesterday’, only the perfective past –the pretérito indefinido– seems to be allowed, as in (7b), while the pretérito perfecto compuesto in (7a) is said to be ungrammatical. that is why the spanish perfect has been defined as a hodiernal (i.e., relating to the present day) past marker, a temporal distinction that has been shown to be at play with specific dedicated markers in some other languages and language families (e.g., comrie 1976, dahl 1985). there are other properties of past-referring temporal adverbials that seem to play a role in their compatibility with perfect markers. for example, some work in english has provided indications that deictic adverbials (i.e., adverbials whose reference is calculated with respect to the speaker’s time/space center of reference) behave differently with respect to their (in)compatibility with the english present perfect (e.g., hitzeman 1995). different from (1), the present perfect seems to be able to combine with deictic past-time referring adverbials that include speech time (s), such as this afternoon, as in (8): (8) chris has left / left york this afternoon. this sentence is acceptable with the present perfect even if this deictic adverb could technically be temporally locating the event of chris’ leaving at the exact same time that the adverbial at three o’clock, as in (1), a sentence that is considered ungrammatical with the present perfect. the main account for this pattern is that deictic adverbs target a relation in which reference time (r) is situated at speech time (s), and therefore do not create an incompatibility with the temporoaspectual configuration of the perfect. however, the role of deixis in the compatibility of the present perfect with past-referring temporal adverbials has not been systematically tested either in english or in spanish. proceedings of elm 2: 61-71, 2023 martı́n fuchs and martijn van der klis: crosslinguistic differences on the present perfect puzzle. 63 https://doi.org/10.3765/elm https://www.elm-conference.net/ additionally, the example in (8) presents an adverbial that is both deictic and hodiernal, but it is also possible that these properties play each an independent role. parallel corpus work (van der klis et al. 2022) actually indicates that standard written peninsular spanish not only allows its perfect to combine with temporal adverbials that locate the event within the day of utterance, but also with deictic adverbials such as este mes ‘this month’, which locate the event within an interval that is calculated from speech time and includes the day of utterance, but can also temporally locate the event before that day (e.g., an event that occurred this month might have happened twelve days ago). this difference in the temporal extension of deictic adverbials (i.e., whether the time span they define includes or is included in the day of utterance) has not been systematically explored either. in the following sections, we present an experimental study on the acceptability of different past-time referring adverbials with the perfect and the past markers of english, spanish, and dutch –with the expectation that the latter works as a control language, since we predict its perfect to have no restrictions to locate events in the past.3 our main goal is to clarify the relationship between the temporal and referential properties expressed through adverbials and the use of perfect and past markers crosslinguistically. to this end, we conducted an acceptability judgment task on the use of perfect and past markers in these three languages to express differently temporally located past events. we consider a twofold distinction of past-referring temporal adverbials: temporal proximity and deixis. first, (6) and (7) indicate variation between adverbials related to the day of utterance and those that are not, showing the relevance of temporal proximity. second, (1) and (8) drive a distinction between deictic and non-deictic adverbials. we describe the specifics of the experimental methodology in section 2, we present the results in section 3, and we discuss them and advance a general conclusion in section 4. 2. methods. 2.1. materials. we investigate uk english, peninsular spanish, and netherlandic dutch4 use of perfect and past markers in combination with different temporal adverbials that we distinguish crossing two independent variables described directly below: • temporal proximity (t), which in +t cases refers to adverbials that are related to the day of utterance by being included in it (e.g., this morning), overlapping with it (e.g., today) or including it (e.g., this month), and conversely, in -t cases designates adverbials such as last month, which do not include or are included in the day of utterance. • deixis (d). this variable refers in +d cases to adverbials whose temporal reference is deictic in nature. for example, to place an adverb such as yesterday on the timeline, we need information about the speaker’s temporal location at reference time. conversely, -d adverbials, such as in november, can be more easily placed on the timeline independently from the speaker’s center of reference.5 3 interestingly, the dutch vtt only presents no restrictions when it deals with events in dialogue or single sentences. to refer to states in the past, or to express narrative discourse (i.e., sequences of events in the past), dutch also needs to resort to its past marker, the ovt (e.g., le bruyn et al. 2019). 4 we decided to test these european varieties since the parallel corpus work that identified some of the constraints that seem to be at work in the distributions of perfect and past markers in these languages (e.g., le bruyn et al. 2019, van der klis et al. 2022) based its claims on originals and translations in these dialectal varieties. 5 of course we claim this ‘independence’ of -d adverbs only in relation to the +d cases, since an adverb such as in november also usually refers to either the previous or the following november with respect to speech time. proceedings of elm 2: 61-71, 2023 martı́n fuchs and martijn van der klis: crosslinguistic differences on the present perfect puzzle. 64 https://doi.org/10.3765/elm https://www.elm-conference.net/ each of these variables has two levels, so that the experimental design, including the grammatical marker as an independent variable, consists of 8 conditions (2x2x2). we created 8 different contexts, for a total of 64 experimental stimuli, presented in appendix a (only in english for reasons of space). we organized these stimuli in a latin square design, so that each participant only saw one experimental condition per context. we additionally created 96 fillers of three different categories: (i) cases that convey either an experiential or a pluractional eventive meaning, with adverbs such as once or twice, which in the three languages are usually expressed with the perfect instead of the past (e.g., de swart, forthcoming); (ii) cases with wh-questions and temporal sluices, which are assumed to be only acceptable with the past in the three languages (see tellings & fuchs, in prep., for reporting of these results); and (iii) unrelated fillers on the use of present and progressive markers to express futurate readings. each stimulus was displayed separately and was accompanied by an introductory context. all experimental sentences convey an achievement to control for lexical aspect (since achievements have no temporal extension, and can be placed directly on the timeline, unlike accomplishments or activities). an example item in english is shown below, with its introductory context in (9) and the different possible experimental conditions in table 1: (9) peter and theresa are planning to go to a concert next weekend. peter offers to go get the tickets later today, but theresa tells him: condition marker adverbial continuation perfect +t, +d i have purchased mine this morning it was cheaper that way. perfect +t, -d at midnight perfect -t, +d last month perfect -t, -d in november past +t, +d i purchased mine this morning past +t, -d at midnight past -t, +d last month past -t, -d in november table 1: experimental conditions, illustrated with an example, crossing three independent variables: grammatical marker, temporal proximity (t) and deixis (d), with two levels each. 2.2. procedure. previous work on this topic has mostly relied on introspection (e.g., klein 1992) or production data such as corpora (e.g., van der klis et al. 2022) or force-choice tasks (e.g., schwenter 1994), where results are binary (presence or absence of a specific marker). moreover, corpus data of specific marker-adverb combinations can be relatively scarce. since large-scale judgment data can provide a more nuanced perspective than corpus data (kepser & reis 2005, francis 2022), not only showing the best form in a given context, but also indicating preferences between forms, we decided to run an online acceptability judgment task in which participants rated individually presented sentences in a 5-point likert scale. they also had to answer yes-no comprehension questions, which followed 75% of the items, included to check that participants were paying attention to the experimental stimuli that they had to rate. each participant saw a total of 8 experimental stimuli and 16 fillers. the allotted time to complete the task was of 15 minutes, with proceedings of elm 2: 61-71, 2023 martı́n fuchs and martijn van der klis: crosslinguistic differences on the present perfect puzzle. 65 https://doi.org/10.3765/elm https://www.elm-conference.net/ an additional 5 minutes to read and sign a consent form and fill in some basic demographic information. participants were compensated with 4 euros for completing the task.6 2.3. participants. 160 subjects per language completed the task, allowing to us to obtain, when combined, 20 full datasets per language under examination. english and spanish participants were recruited through amazon mturk, where we were able to control for their ip addresses to make sure that they were from the united kingdom and spain, respectively. they had to indicate the region where they lived, and, by self-report, they were from all regions of the united kingdom and spain, but mostly from the greater london and greater madrid areas. in the case of dutch, we recruited participants through neerlandistiek.nl, a website dedicated to dutch language and culture. participants also came from all regions of the netherlands, but were mostly from the utrecht region. 3. results. all participants performed above 75% accuracy in the comprehension questions, so no participant was excluded from data analysis. the included fillers worked as we had expected (high ratings for the perfect and low ratings for the past across languages in filler condition (i) / low ratings for the perfect and high ratings for the past in wh-questions and sluices, in condition (ii)), providing additional support that participants were sensitive and attentive to the task. we do not find significant differences in the ratings across geographical regions in any of the languages, so we report the data in full per language on the remainder of the paper.7 mean acceptability scores per experimental condition in each of the languages under study are reported in table 2: type of adverbial marker english spanish dutch +t, +d (this morning) perfect 4.03 4.05 4.28 past 4.42 4.31 3.37 +t, -d (at midnight) perfect 3.34 4.33 3.78 past 4.33 4.03 3.14 -t, +d (last month) perfect 3.42 3.14 4.37 past 4.51 4.53 3.58 -t, -d (in november) perfect 3.44 3.21 4.07 past 4.53 4.53 3.19 table 2: mean acceptability scores by experimental condition in each language. since we are interested in the distribution of these markers and their compatibility with different kind of adverbials in each of these languages, we performed separate statistical analysis per language. we subjected the data to linear mixed-effect analysis, which were performed with random intercepts for subject and item, and fixed effects for the interaction of grammatical marker, temporal proximity and deixis. 6 the study was approved by the faculty ethics assessment committee of the faculty of humanities (fetc-h) of utrecht university (reference number: 20-249-03). 7 this was particularly surprising in the case of spain, since previous reports indicate a wider use of the past in some of the northern regions (e.g., azpiazu 2013, 2015). we consider that the lack of variability in our study is probably due to the nature of the task, which allows to rate both markers as (un)acceptable. however, we recognize that further exploration with a more controlled recruitment process could provide results where that dialectal variation within peninsular spanish is observed. proceedings of elm 2: 61-71, 2023 martı́n fuchs and martijn van der klis: crosslinguistic differences on the present perfect puzzle. 66 https://doi.org/10.3765/elm https://www.elm-conference.net/ in dutch, we find a main effect of marker (χ2(2) = 32.117; p < .001), favoring the perfect over the past across all conditions (β = 0.8031; p < .001). figure 1 shows the means across conditions in this language. figure 1: mean acceptability scores by experimental condition in dutch (*** = p < .001; ** = p < .01, * = p < .05; ns = not significant). spanish shows a significant interaction of temporal proximity*marker (χ2(1) = 47.12; p < .001), with no effect of deixis. in the -t condition, there is a main effect of marker (χ2(1) = 57.07; p < .001), favoring the past over the perfect (β = 1.353; p < .001), but in the +t condition, there is no significant effect of marker (χ2(1) = 0.016; p = .90). a summary in terms of means per experimental condition for spanish is shown in figure 2. figure 2: mean acceptability scores by experimental condition in spanish (*** = p < .001; ** = p < .01, * = p < .05; ns = not significant). finally, in english, there is a significant effect of temporal proximity*deixis*marker (χ2(2) = 6.373; p < .05), and a main effect of marker, favoring the past over the perfect in all conditions proceedings of elm 2: 61-71, 2023 martı́n fuchs and martijn van der klis: crosslinguistic differences on the present perfect puzzle. 67 https://doi.org/10.3765/elm https://www.elm-conference.net/ (β= 1.0875; p < .001). the interaction effect arises because there is less of a categorical difference in the +t,+d condition, but a post-hoc test with tukey correction still shows the effect of marker (β= 0.394; p = .035). figure 3 presents a bar graph with the means per condition in english. figure 3: mean acceptability scores by experimental condition in english (*** = p < .001; ** = p < .01, * = p < .05; ns = not significant). an interesting, more nuanced result is also revealed in english when subdividing +t,+d adverbials by whether the adverb includes the day of utterance or is included in it. in those cases, we find a significant marker effect in the first case(χ2(1) = 6.7711; p <.01) favoring the past over the perfect (β= 0.5931; p < .001), but the effect disappears in adverbs included in the day of utterance (χ2(1) = 0.5942; p = .4408; perfect mean = 4.25; past mean = 4.38), where both markers produce ratings not significantly different. a bar graph showing this contrast is presented in figure 4. figure 4: mean acceptability scores in english for conditions that include adverbs that are +temporal proximity and + deixis, distinguished by whether the adverb includes the day of utterance (e.g., this month) or is included in it (e.g., this morning) (*** = p < .001; ** = p < .01, * = p < .05; ns = not significant). proceedings of elm 2: 61-71, 2023 martı́n fuchs and martijn van der klis: crosslinguistic differences on the present perfect puzzle. 68 https://doi.org/10.3765/elm https://www.elm-conference.net/ in summary, dutch speakers prefer the vtt –its perfect marker– over the ovt –its past marker– across the board. spanish speakers accept the perfect when the adverb is linked to the present, but there is no preference for the pretérito perfecto compuesto in the +t condition: the pretérito indefinido receives similar ratings. english speakers prefer the simple past in all conditions but they accept the present perfect with deictic hodiernal adverbials, especially when the adverb is included in the day of utterance. 4. discussion and general conclusion. this study had the objective of experimentally testing the acceptability of perfect and past markers in combination with various past-time referring adverbials in english, spanish, and dutch, with the ultimate goal of providing a more thorough and crosslinguistically valid account of the present perfect puzzle (klein 1992). based on previous claims in the literature, we assumed that the dimensions of variability in the temporal and referential properties of adverbials that could have an effect on the acceptability of these markers were whether the adverb indicates a time span that is included in the day of utterance (i.e., hodiernality) and whether the time span is calculated from the speaker’s center of reference (i.e., deixis). firstly, the results from our acceptability judgment task confirm the patterns previously reported in the literature in these languages (e.g., van der klis et al. 2022 for dutch; schwenter 1994 for spanish; klein 1992, hitzeman 1995 for english). dutch allows its perfect to combine with any kind of past-referring adverbial, spanish only allows the perfect form to appear when adverbials are hodiernal, and english prefers its past marker to make reference to past events, but allows its perfect to appear when the adverbs that it combines with are both deictic, and indicate a time span included in the day of utterance. additionally, our experimental data refines these generalizations with respect to the distribution of perfect and past markers across languages. the behavioral results not only provide wider empirical support for previous claims, but also deepen our understanding about the effect that temporal and referential properties expressed by past-referring adverbials might have on the crosslinguistic distribution of these markers, revealing some patterns previously undescribed. in the case of spanish, we did not only show that the pretérito perfecto compuesto is allowed to express events that adverbials situate within the day of utterance (e.g., this morning), but also that these events can be situated further away in time, as long as the reference time is kept at speech time so that the time span that the adverb indicates is calculated from the deictic center (e.g., this month). moreover, and contra previous descriptions in the literature (e.g., schwenter 1994, azpiazu 2013), our experimental data also revealed that the spanish perfect may be the preferred form to express these hodiernal events, but the pretérito indefinido –the spanish perfective past– is also acceptable in these contexts. the english experimental data also provides an additional insight. the english present perfect is allowed to combine with adverbials that are both deictic and that signal an interval which is temporally close to speech time. however, to properly account for the patterns in the data, we need to make a distinction between ‘proper hodiernality’ and ‘extended hodiernality’, since english only allows its perfect to combine with deictic adverbials that are ‘properly’ hodiernal –that is, adverbials that indicate a time span included in the day of utterance–, and rejects the use of the perfect when the deictic adverbial indicates an interval that includes the day of utterance. this distinction had not been previously reported, and might be relevant for understanding the distribution of perfect and past markers more widely. all in all, we conclude that, aside from providing results that specify a finer grained picture of the present perfect puzzle from a crosslinguistic perspective, this study functions as an proceedings of elm 2: 61-71, 2023 martı́n fuchs and martijn van der klis: crosslinguistic differences on the present perfect puzzle. 69 https://doi.org/10.3765/elm https://www.elm-conference.net/ illustration of how behavioral methods in experimental linguistics can be brought to bear on classical theoretical problems in semantics and pragmatics. references azpiazu, susana. 2013. antepresente y pretérito en el español peninsular: revisión de la norma a partir de evidencias empíricas. anuario de estudios filológicos 36, 19-31. azpiazu, susana. 2015. la variación antepresente/pretérito en dos áreas del español peninsular. verba 42, 269-292. https://doi.org/10.15304/verba.42.1371 comrie, bernard. 1976. aspect. an introduction to the study of verbal aspect and related problems. cambridge: cambridge university press. dahl, östen. 1985. tense and aspect systems. oxford: basil blackwell. francis, elaine j. 2022. gradient acceptability and linguistic theory. oxford: oup. fuchs, martín & paz gonzález. 2022. perfect-perfective variation across spanish dialects: a parallel-corpus study. languages 7(3), 166. https://doi.org/10.3390/languages7030166 gonzález, paz, margarita jara yupanqui, & carmen kleinherenbrink. 2019. the microvariation of the spanish perfect in three varieties. isogloss 4: 115–33. https://doi.org/10.5565/rev/isogloss.60 harris, martin. 1982. the ‘past simple’ and the present perfect in romance. in nigel vincent and martin harris (eds.), studies in the romance verb: essays offered to joe cremona on the occasion of his 60th birthday. 42-70. london: croom helm. hitzeman, janet. 1995. a reichenbachian account of the interaction of the present perfect with temporal adverbials. proceedings of the north east linguistic society (nels) 25 (1), 17. howe, chad. 2006. cross-dialectal features of the spanish present perfect: a typological analysis of form and function. columbus, oh: the ohio state university dissertation. kepser, stephan & marga reis. 2005. linguistic evidence: empirical, theoretical and computational perspectives. berlin, new york: de gruyter mouton. klein, wolfgang. 1992. the present perfect puzzle. language 68, 525-552. https://doi.org/10.2307/415793 klis, martijn van der, bert le bruyn and henriëtte de swart. 2022. a multilingual corpus study of the competition between past and perfect in narrative discourse. journal of linguistics 58(2), 423-457. https://doi.org/10.1017/s0022226721000244 le bruyn, bert, martijn van der klis & henriëtte de swart. 2019. the perfect in dialogue: evidence from dutch. linguistics in the netherlands 36 (1), 162-175. https://doi.org/10.1075/avt.00030.bru musan, renate. 2002. the german perfect: its semantic composition and its interaction with temporal adverbials. dordrecht: kluwer. reichenbach, hans. 1947. elements of symbolic logic. london: mcmillan schaden, gehrard. 2009. present perfects compete. linguistics and philosophy 32(2): 115-141. https://doi.org/10.1007/s10988-009-9056-3 schwenter, scott. 1994. the grammaticalization of an anterior in progress: evidence from a castilian spanish dialect. studies in language 18: 71-111. https://doi.org/10.1075/sl.18.1.05sch squartini, mario & pier marco bertinetto. 2000. the simple and compound past in romance languages. in östen dahl (ed.), tense and aspect in the languages of europe. 403-440. berlin: de gruyter. swart, henriëtte de. 2007. a cross-linguistic discourse analysis of the perfect. journal of pragmatics 39, 2273-2307. https://doi.org/10.1016/j.pragma.2006.11.006 proceedings of elm 2: 61-71, 2023 martı́n fuchs and martijn van der klis: crosslinguistic differences on the present perfect puzzle. 70 https://doi.org/10.3765/elm https://www.elm-conference.net/ swart, henriëtte de. forthcoming. perfect variations in romance. isogloss. tellings, jos & martín fuchs. in prep. sluicing and temporal definiteness. vet, co. 1980. temps, aspects et adverbes de temps en français contemporain. genève: droz. vet, co. 1992. le passé composé: contextes d’emploi et interprétation. cahiers de praxématique 19, 37–59. appendix a: experimental stimuli 1. patrick and carl are roommates who want to spend the summer abroad, but they are not sure where to go. the application for international internships is due tomorrow. when patrick sees carl at home, carl tells him: i’m very happy, i signed / have signed a rental contract for a new york city apartment (at two in the afternoon / this evening / on monday / last night). 2. andrew is looking for his earplugs, but he cannot find them. he’s on the phone with his friend paul, who tells him: that’s frustrating, i know how you feel. mine were lost, but i have found / found them (at noon / this week / on friday / last weekend). 3. sandra and christine are planning their wedding. there’s little time left before the ceremony, but there are still many things that need to be done. while having dinner, christine tells sandra: oh, by the way, i picked up / have picked up the invitations (at four / this afternoon / on sunday / the day before yesterday), so that’s one less thing on our list. 4. sandra used to have an old tv that her friend nick would always complain about whenever they watched tv at her place. this time, when nick comes to visit, sandra tells him about a new appliance store in town, and announces: i bought / have bought a new tv (at nine am / today / on tuesday / last week) at their grand opening sale. 5. linda and frank are meeting for dinner at the italian restaurant where linda also works. when frank gets there and looks at the menu, linda tells him: i tried / have tried the lasagna (at lunch / this month / on wednesday / yesterday). it is delicious. 6. peter and theresa are planning to go to a concert next weekend. peter offers to go get the tickets later today, but theresa tells him: i purchased / have purchased mine (at midnight / this morning / in november / last month), during the pre-sale. it was cheaper that way. 7. nicholas and josephine are participating in a chess tournament. they are discussing their next opponents. josephine is about to play against william, so nicholas tells her: he defeated / has defeated me (at three / this year / in 2019 / last year), so you’d better be careful. 8. sam has to take an oral german examination this afternoon, and he is a bit nervous about it. his friend alex tells him: laura passed / has passed the exam (at ten o’clock in the morning / this semester / in january / last winter), so you should talk to her. proceedings of elm 2: 61-71, 2023 martı́n fuchs and martijn van der klis: crosslinguistic differences on the present perfect puzzle. 71 https://doi.org/10.3765/elm https://www.elm-conference.net/ toward an accommodation account of deaccenting under nonidentity jeffrey geiger & ming xiang* abstract. two competing models attempt to explain the deaccentuation of antecedentnonidentical discourse-inferable material (e.g., bach wrote many pieces for viola. he must have loved string instruments). one uses a single grammatical constraint to license deaccenting for identical and nonidentical material. the second licenses deaccenting grammatically only for identical constituents, whereas deaccented nonidentical material requires accommodation of an alternative antecedent. in three experiments, we tested listeners’ preferences for accentuation or deaccentuation on nonidentical inferable material in out-of-the-blue contexts, supportive discourse contexts, and in the presence of the presupposition trigger too. the results indicate that listeners by default prefer for inferable material to be accented, but that this preference can be mitigated or even reversed with the help of manipulations in the broader discourse context. by contrast, listeners reliably preferred for repeated material to be deaccented. we argue that these results are more compatible with the accommodation model of deaccenting licensing, which allows for differential licensing of deaccentuation on inferable versus repeated constituents and provides a principled account of the sensitivity of accentuation preferences on inferable material to broader contextual manipulations. keywords. deaccenting; prosody; information structure; givenness; inference; identity; accommodation; entailment 1. introduction. it has long been recognized that there is a close connection between a constituent’s information status within a discourse and its prosodic realization. canonically, contentful discourse-new constituents are realized with a high or rising pitch accent (chafe 1974, pierrehumbert & hirschberg 1990). in english, this is most noticeable on the most embedded constituent, which receives a nuclear pitch accent, impressionistically the most prominent accent in the clause (chomsky & halle 1968, selkirk 1984, büring 2016). for example, in (1), coffee exhibits a nuclear pitch accent, denoted by small caps. (1) the caterer asked if there were any drinks to avoid, and i said i don’t like coffee. in contrast to new constituents, constituents with an identical correlate in a local linguistic antecedent tend not to exhibit a high pitch accent. such constituents are exempt from the default rules of stress assignment via an operation called deaccenting or stress shift, and are realized as prosodically reduced, according to both listener judgments and a phonetic correlates of emphasis such as intensity, f0, and duration (pierrehumbert & hirschberg 1990, chodroff & cole 2019, geiger & xiang 2019). although the final coffee in (2) is in a structural position to receive a nuclear pitch accent, it is deaccented by virtue of the instance of coffee in the previous clause, and the nuclear accent falls instead on like. (2) the caterer asked if they should supply coffee, and i said i don’t like coffee. *this work was supported by national science foundation grant no. bcs-1827404. authors: jeffrey geiger, university of chicago (jeffrey.geiger@pomona.edu) & ming xiang, university of chicago (mxiang@uchicago.edu). proceedings of elm 1: 172-183, 2021 c©2021 jeffrey geiger and ming xiang published by the lsa with permission of the author(s) under a cc by license. 172 https://doi.org/10.3765/elm https://www.elm-conference.net/ this mapping between givenness and deaccentuation exhibits a number of complications, however. one is that constituents are typically deaccented only when they are in a structural position isomorphic to that of their antecedent (cf. #mary saw john, then kim saw mary; tancredi 1992, schwarzschild 1999). this is not the primary focus of the current paper, but is an important consideration in the grammar of deaccentuation. a second complication is that the literature recognizes a number of cases in which a constituent can felicitously be deaccented despite not having an identical correlate in a local linguistic antecedent. this includes constituents that corefer with an antecedent, as in (3); constituents that are entailed by an antecedent constituent, modulo existential closure, as in (4); constituents whose meanings are made salient by synthesizing information from the antecedent with broader world knowledge, as in (5); and constituents whose meanings are made salient by the nonlinguistic context, as in (6). (3) a: did you see dr. cremer to get your root canal? b: don’t remind me. i’d like to strangle the butcher. (büring 2007) (4) bach wrote many pieces for viola. he must have loved string instruments. (van deemter 1999) (5) first john called mary a republican, and then she insulted him. (lakoff 1968) (6) [hearer cocks their head to one side as if listening for a faint or distant noise.] speaker: i heard it, too. (rochemont 1986) the diverse configurations that give rise to deaccentuation, including under identity, as in (2), and under nonidentity, as in (3) through (6), raise the question of how to characterize the set of deaccentable material, and what combination of grammatical and extragrammatical mechanisms generates deaccented structures. two classes of solution have emerged in the literature. the first approach, which for convenience we refer to as the grammatical model of deaccentuation licensing, concentrates on developing a single condition under which deaccenting is licensed; this condition accounts for the deaccentuation of both antecedent-identical and -nonidentical material. the fundamental insight of the grammatical approach is that givenness is the underlying property that characterizes deaccentable material, and that constituents with identical antecedents merely constitute a trivial subset of given material. some grammatical approaches posit essentially a one-to-one mapping between givenness and deaccentuation, with the challenge lying in determining what material counts as given. for instance, rochemont (1986) exempts constituents from focus marking when they are c-construable, a flexible notion of givenness that allows deaccentuation of material entailed by an antecedent, salient in the nonlinguistic context, or treated as background information by the speaker and hearer. van deemter (1994, 1999) suggests that a constituent can be deaccented when it is object-given, meaning that it corefers with an antecedent, or concept-given, when it is entailed by an antecedent. baumann and riester (2012) make a similar appeal in their reflex annotation scheme, where a proceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 173 https://doi.org/10.3765/elm https://www.elm-conference.net/ constituent can be “referentially” (r-)given if it corefers with an antecedent, or “lexically” (l-)given if it is identical to or entailed by an antecedent via synonymy or a superset relation. other grammatical approaches interface more directly with focus theory, and aim to simultaneously capture the structural constraints on deaccentuation and the mapping between notional givenness and deaccenting. one such account comes from rooth (1992). under this model, the use of any focus-marked structure introduces a presupposition that there is a focus alternative to the structure that is contextually available. this requirement is evaluated along semantic lines, meaning that non-f-marked (deaccented) material is felicitous when it is entailed by an antecedent, whether or not the constituents are string-identical (e.g., string instruments with an antecedent of viola in (4)). in another influential model, schwarzschild (1999) requires that every substructure of a sentence be given, meaning informally that the context supplies a meaning that entails the constituent with its f-marked components replaced by variables. as in rooth’s account, the requirement that deaccented material be merely entailed by the antecedent allows for the deaccenting of material that does not necessarily have a string-identical correlate in the antecedent. crucially, in both accounts, antecedent-identical constituents are deaccentable because they are trivially entailed by their identical antecedent. in contrast to the grammatical approach, the accommodation model evokes two separate mechanisms for generating deaccentuation on antecedent-identical constituents versus nonidentical constituents. according to the accommodation model, the grammatical property characterizing deaccentable material is string identity with an antecedent, not givenness. thus, deaccentuation is licensed on antecedent-identical constituents using this grammatical mechanism. by contrast, deaccentuation of antecedent-nonidentical material, as in (3) through (6), is strictly ungrammatical. however, such instances of deaccentuation can be marked acceptable via a second, extragrammatical mechanism. if it is reasonable to do so, an alternative antecedent can be accommodated that contains an identical correlate to the deaccented material; the deaccenting is then treated as acceptable on the basis of identity with this accommodated antecedent. in an early accommodation approach, tancredi (1992) proposes that deaccentuation requires prior “instantiation” in the antecedent (string identity), but that the linguistic context can be augmented via “pragmatic incrementation” to include implied alternative utterances that would license observed instances of deaccentuation under nonidentity. fox (2000) similarly characterizes non-fmarked (deaccented) material without an identical correlate in the antecedent as accommodationseeking material, meaning its use triggers the accommodation of an alternative linguistic structure that would have licensed deaccentuation according to the grammatical identity requirement. finally, while wagner (2012) does not concentrate on an accommodation operation, he makes reference to how such a mechanism could underlie cases of deaccentuation under nonidentity. despite sustained interest in the connection between givenness and deaccentuation, there has been little systematic quantitative research on the prosodic properties of inferable (discourseaccessible) material, such as constituents entailed by, but not identical to, an antecedent. two recent exceptions include chodroff & cole (2019) and geiger & xiang (2019), who examined the production of new, given, and inferable nouns and verbs, respectively. each found that discourseaccessible constituents tend to be realized more similarly to discourse-new than given material, casting doubt on the assumption in both the grammatical and accommodation literature that deacproceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 174 https://doi.org/10.3765/elm https://www.elm-conference.net/ centing under nonidentity is straightforward. the present paper builds on these production studies by examining listener assessments of prosodic naturalness for new, repeated, and inferable verbs in perception. whereas it is possible that deaccentuation is optional in production, or that speakers in the prior production studies may not have been aware of the relevant inferencing relations making the critical constituents inferable, listeners in a perception study must contend with the fact that the speaker chose to accent or deaccent a particular constituent. this may lead to more robust conclusions about whether deaccentuation is licensed on inferable constituents, and about what licensing mechanisms underlie deaccentuation for antecedent-identical and -nonidentical material. 2. experiment 1. experiment 1 assessed the naturalness of accentuation and deaccentuation on verbs of three different levels of information status: new, inferable, and repeated (given). the goal was to compare listeners’ prosodic preferences for inferable verbs to those for new and repeated verbs, which canonically should sound more natural when they are accented and deaccented, respectively. crucially, the grammatical licensing model predicts that highly inferable and repeated verbs should exhibit roughly the same behavior, since they are subject to the same constraint on deaccenting, while the accommodation model leaves room for differential behavior between inferable and repeated verbs, since they are treated by separate deaccenting licensing mechanisms. 2.1. design and materials. the critical materials in the experiment were audio-recorded sentences of the form svo and svo. the full experiment had a 3 (verb status) × 2 (object status) × 2 (accent) design. within each item, the second svo string remained constant, while the first svo clause varied to determine the discourse status of the second-clause constituents. all subjects and objects were proper names, and the second-clause subject was always different from the firstclause subject. table 1 outlines the conditions discussed in this paper, which include the full verb status and accent manipulations, but only one level for object status. accent verb status recording new gabriel punished amy, and nan surprised amy. accented inferable ethan astounded amy, and nan surprised amy. repeated benjamin surprised amy, and nan surprised amy. new gabriel punished amy, and nan surprised amy. deaccented inferable ethan astounded amy, and nan surprised amy. repeated benjamin surprised amy, and nan surprised amy. table 1: sample experiment 1 stimuli. small caps: nuclear accent. italics: deaccentuation. the second-clause verb could be new, repeated, or inferable. new verbs did not stand in a clear inferencing relationship with the first-clause verb, such as surprised following punished. repeated verbs were identical to the first-clause verb, as in surprised following surprised. inferable verbs stood in a nonidentical inferencing relation to the antecedent verb, such that they were discourse-accessible, but not given. two such inferencing relations were tested, each in one half of the experimental items: entailment (e.g., astounded-surprised), and more informal semantic relatedness that might support the conclusion that the second verb was accessible despite not being entailed by the first (e.g., charmed-seduced). the meanings of the inferable verbs were rated as highly available in the context of their antecedents in a separate norming study, whereas inferabilproceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 175 https://doi.org/10.3765/elm https://www.elm-conference.net/ ity ratings for the new verbs were low (geiger & xiang 2019).1 for the experiments presented here, the results for the two types of inferencing relations were qualitatively very similar, so we collapsed the two categories and the results are presented together as one inferable condition. the second-clause object could be old or new. old objects were identical to the object of the first clause, while new objects were a different proper name not used anywhere else in the sentence. since an old object should canonically be deaccented and a new object accented, this manipulation determined whether the second-clause verb was in nuclear (preceding an old object) or prenuclear (preceding a new object) position. ratings did not reliably vary as a function of either verb status or accent in the new-object conditions, so only the results from the old-object conditions are presented and discussed in this paper. finally, the second-clause verb could be accented or deaccented. two naive participants, one female and one male, read aloud svo and svo sentences featuring the verb status × object status manipulation described above (for more information, see geiger & xiang 2019). in this paradigm, new verbs canonically should be accented, while repeated verbs should be deaccented. thus, to generate sentences with the appropriate second-clause verb accent, recordings of second clauses (including and) with typical accentuation (read as new) or typical deaccentuation (read as repeated) were cross-spliced with the appropriate first-clause recordings to generate sentences with each verb status/object status combination. 2.2. procedure. participants were gathered via amazon mechanical turk with the stipulations that they be at least 18 years old, native speakers of english, and using a computer in the united states. upon accepting the task, participants were redirected to the experiment, conducted in ibex farm (drummond 2020). they provided informed consent, verified that the experiment site played sound at an appropriate volume on their system, filled out a demographic survey, and completed two practice trials to familiarize them with the setup and response protocol. in each critical trial, participants first viewed a preview screen that displayed, as text, the sentence they would be rating. participants then pressed any key to advance from the preview screen to the test screen. upon advancing, the audio file played automatically over the participant’s speakers. participants were prompted on this screen to rate how natural they found the “melody” or “tune” of the sentence, where 1 represented the least natural and 7 represented the most natural. the experiment consisted of 24 critical trials as well as 10 filler trials featuring 5 canonical and 5 noncanonical productions of english prosodic contours (e.g., declaratives, polar questions, lists). the experiment took approximately 8 minutes to complete. 2.3. participants. 142 participants (67 female, mean age 36.6 years) took part in the experiment. the data from one participant was excluded because they failed to identify themselves as a native speaker of english, and the data from a further five participants was excluded due to inattention (mean response time less than 1000 milliseconds). 2.4. results and analysis. the results of experiment 1 are shown in figure 1. a visual inspection of the plot shows that as expected, participants rated sentences with new second-clause 1sample prompt: suppose you know that alice astounded billy. how likely do you think it is that alice surprised billy? the suppose sentence contained the first-clause antecedent verbs, conditioning either a new or inferable relation, while the rating question contained the target second-clause verb. mean inferability score for inferable verbs: 6.14/7. mean score for new verbs: 2.14/7. proceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 176 https://doi.org/10.3765/elm https://www.elm-conference.net/ verbs as sounding more natural if that verb was accented than if it was deaccented. conversely, sentences with repeated verbs received higher ratings when the verb was deaccented. crucially, the results for inferable verbs resemble those for discourse-new verbs, with participants preferring accentuation over deaccentuation. figure 1: experiment 1 results. error bars: 95% confidence interval. to explore the results in further detail, they were fit to a linear mixed-effects regression model with an interaction of verb status and accent, main effects of verb status and accent, and random intercepts for participant and item.2 there was a significant interaction between verb status and accent (p<.001), and the main effects of verb status (p<.001) and accent (p<.05) were significant. paired comparisons were carried out using estimated marginal means to test for an accentuation preference within each verb status level (i.e., new, inference, repeated). these paired comparisons indicated that the ratings of accentedversus deaccented-verb sentences differed significantly within all three levels of verb status (p’s<.001). 2.5. discussion. the paired comparisons indicated that participants significantly preferred sentences where new verbs were accented rather than deaccented, and they preferred sentences where repeated verbs were deaccented rather than accented. these results are unsurprising, as they conform to the canonical accent patterns for new and given material in nuclear position. crucially, participants significantly preferred sentences in which inferable verbs were accented over those in which they were deaccented. this contrasts somewhat with the predictions of both the grammatical and the accommodation accounts of deaccenting under nonidentity. both models are meant to generate deaccentuation of inferable material, including entailed constituents. by contrast, the experimental results indicate that participants tend to treat inferable constituents as though they are discourse-new, with accentuation reliably preferred over deaccentuation. despite the mismatch between the experimental results and the prediction made by both accounts that deaccentuation of inferable material should be acceptable, the results are less problematic for the accommodation account than for the grammatical account. the grammatical model 2lmer model specification: response ∼ verbstatus * accent + (1|participant) + (1|item). proceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 177 https://doi.org/10.3765/elm https://www.elm-conference.net/ uses one mechanism to generate deaccenting, with deaccenting under identity merely a subset of the cases in which the givenness requirement holds. thus, it is not clear that this model can predict such a wide gap in (de)accentuation preferences for inferable versus repeated material. by contrast, the accommodation model uses two mechanisms to generate deaccenting for inferable versus repeated material. the preference for accentuation on inferable constituents is somewhat at odds with the intended prediction of the model. however, since deaccentuation of inferable material is generated by an extragrammatical accommodation mechanism, whereas deaccentuation of repeated material is triggered by identity with an antecedent, the differential accentuation preferences for inferable and repeated material are not as clearly problematic for this model. 3. experiment 2. experiments 2 and 3 focus on further exploring the feasibility of the accommodation account. this model proposes that deaccenting under nonidentity relies on the extragrammatical accommodation of an alternative antecedent that would license deaccenting under identity. the pragmatic nature of this mechanism suggests that it should be sensitive to manipulations in the broader discourse context that might encourage or discourage such accommodation. whereas experiment 1 presented the critical sentences out of the blue, experiment 2 explored the role of the broader discourse by introducing a supportive context beyond the critical sentence. if the broader context supports parallel readings for inferable verbs and their antecedents, participants might find it more natural to accommodate the necessary alternative antecedent according to the accommodation model, increasing ratings for tokens with deaccented inferable verbs. 3.1. design, materials, and procedure. the design, materials, and procedure were identical to those for experiment 1, with one modification. in experiment 2, the preview screen on each trial now displayed a context sentence instead of a text representation of the sentence the participant would hear. participants were told that the sentence on the warning screen represented a scenario, and that the sentence they heard represented someone talking about the scenario. the purpose of the context sentence was to promote a “pragmatically identical” reading of the inferable verb and its antecedent. for instance, for the critical inferable-verb sentence veronica bullied roy, and kendall intimidated roy, the context sentence was as they did every year, the teachers worried about how the students would interact with each other on the first day of high school. the purpose of the context sentence here is to suggest that bullying and intimidation represent situationally comparable actions in the context – both are clearly negative interactions, and they do not obviously contrast with one another. 3.2. participants. 144 participants (53 female, mean age 33.5) took part in experiment 2. the data from two participants was excluded from analysis because they failed to self-identify as native english speakers, while the data from a further seven participants was excluded due to inattention. 3.3. results and analysis. the results of experiment 2 are shown in figure 2. as in experiment 1, the plot appears to suggest a reliable preference for accentuation on new verbs and deaccentuation on repeated verbs. accentuation is once again preferred for inferable verbs, but the magnitude of this preference appears to be reduced relative to experiment 1. the initial analysis conducted for experiment 2 was identical to that for experiment 1. the linear mixed-effects regression model showed a significant interaction of verb status and accent (p<.001). the main effect of verb status was significant (p<.001), while the main effect of accent proceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 178 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2: experiment 2 results. error bars: 95% confidence interval. was not (p>.5). paired comparisons using estimated marginal means indicated that accented and deaccented tokens received significantly different ratings within all three levels of verb status (new p<.001, inferable p<.01, repeated p<.001). to compare the magnitude of the preference for accentuation on inferable verbs in each experiment, the inferable-verb data from experiments 1 and 2 was fit to a linear mixed-effects regression model with an interaction of experiment and accent, main effects of experiment and accent, and random intercepts for item and participant.3 the interaction of experiment and accent was significant (p<.05), as were the main effects of experiment (p<.01) and accent (p<.001). 3.4. discussion. as in experiment 1, the paired comparisons for the experiment 2 results indicated an expected preference for accentuation on new verbs and for deaccentuation on repeated verbs. for inferable verbs, there was still a preference for accentuation over deaccentuation. however, the experiment-accent interaction in the second model indicates that the preference for accentuation on inferable verbs was significantly reduced in magnitude in experiment 2 compared to experiment 1. this finding is compatible with the accommodation account, with broad contextual support for an identical reading of inferable verbs and their antecedents contributing to participants’ willingness to accommodate the appropriate antecedent to license deaccenting. by contrast, it would be difficult for the grammatical account to explain such a difference between the two experiments, since the norming study indicated that participants found inferable verbs to be highly available even in the out-of-the-blue contexts of experiment 1, and this type of givenness is the sole determiner of accent according to the grammatical model. 4. experiment 3. like experiment 2, experiment 3 explores how manipulations in the discourse context outside the critical constituents can affect listeners’ judgments of the naturalness of accentuation and deaccentuation on inferable constituents. where experiment 2 explored the effect of adding information supporting deaccentuation to the broad discourse context, experiment 3 tests the effect of adding the presupposition trigger too to the end of each stimulus sentence. 3lmer model specification: response ∼ experiment * accent + (1|item) + (1|participant). proceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 179 https://doi.org/10.3765/elm https://www.elm-conference.net/ on an informal level, too should indicate that the content of the second clause builds in some way on the discourse contribution of the first clause (beaver & clark 2008). similarly to the role played by the supportive context in experiment 2, this might promote readings of the critical sentence where inferable verbs and their antecedents are treated as “situationally identical”, indicating instances of identical or comparable events rather than contrasting or unrelated ones. in turn, this might encourage accommodation of the necessary alternative antecedent to license deaccenting, increasing the ratings for sentences with deaccented inferable verbs. 4.1. design, materials, and procedure. the design, materials, and procedure were identical to experiment 1, except that each critical sentence was augmented with the word too at the end of the second clause (e.g., ethan astounded justin, and nan surprised amy, too). this addition necessitated re-recording the stimuli, as too was not present in the original recordings. the same female speaker who recorded half of the experiment 1 stimuli re-recorded all of the necessary sentences using the production paradigm described in geiger & xiang (2019) with too added to the end of each critical sentence. the clauses of these recordings were then cross-spliced as in experiment 1 to create the same 3 × 2 × 2 design (only old-object trials reported here). 4.2. participants. 140 participants (46 female, mean age 37.0 years) took part in experiment 3. the data from 23 participants was excluded from analysis due to inattention. 4.3. results and analysis. the results of experiment 3 are shown in figure 3. once again, the plot suggests a preference for deaccentuation on repeated verbs. in contrast to the prior results, however, there appears to be no preference between accentuation and deaccentuation on new verbs. finally, for inferable verbs, deaccentuation appears to be preferred for the first time. figure 3: experiment 3 results. error bars: 95% confidence interval. the analysis for experiment 3 was identical to those for the previous experiments. the linear mixed-effects regression model showed a significant interaction of verb status and accent (p<.001). the main effect of accent was significant (p<.001), while the main effect of verb status was not significant (p<.1). paired comparisons indicated that accentedand deaccented-verb trials received significantly different ratings within the inferable and repeated levels for verb status (p’s<.001), proceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 180 https://doi.org/10.3765/elm https://www.elm-conference.net/ but not within the new level (p>.4). 4.4. discussion. the addition of too had a substantial impact on ratings compared to the results for experiment 1. most critically, the previous reliable preference for accentuation on inferable constituents gave way to a reliable preference for deaccentuation. this suggests that too promoted a reading where the inferable-verb second clause builds on the contribution of the first clause, encouraging treatment of the inferable verbs as “given” and, by extension, deaccentuation. interestingly, while the canonical preference for deaccentuation of repeated verbs persisted, there was no longer a reliable preference for accenting new verbs. this may be connected to the impressionistic observation that the rating “floor” in this experiment was higher than in the previous experiments; even relatively “bad” prosodic configurations received fairly high ratings. this suggests that the addition of too had substantial ameliorative effects on naturalness judgments across the board, indicating that the results of this experiment should be interpreted with caution. nevertheless, the results of experiment 3 appear to support the accommodation model, with the addition of too suggesting that inferable verbs should be treated as though they were given, namely, be deaccented. this indicates that, contra the treatment of the accommodation account in the literature, it may not be sufficient for hearers merely to encounter a deaccented inferable verb for it to be marked as acceptable, but that such instances can be found acceptable in the presence of additional mitigating factors, such as presupposition triggers. by contrast, it is once again unclear how the grammatical model would account for the difference in behavior in the inferable conditions of experiments 1 and 3, since the addition of too does not clearly affect the predictions of this model when the verbs in question were already highly inferable in experiment 1. 5. general discussion. considered together, the results of the three experiments are more compatible with the dual-mechanism accommodation account of deaccenting under nonidentity than the single-mechanism grammatical account. both models are intended to generate canonically cited examples of deaccented antecedent-nonidentical material, such as constituents coreferring with or entailed by an antecedent. by contrast, experiment 1 indicated that listeners preferred for highly inferable material, including both entailed verbs and verbs linked to their antecedents by more informal inferencing relations, to be accented as though they were discourse-new. while this finding is at odds with the goal that both models should generate deaccenting under nonidentity, it is more problematic for the grammatical account than the accommodation account. the reason for this is the substantial difference between the preference for accentuation on inferable constituents and the preference for deaccentuation on repeated constituents. the grammatical model uses a single mechanism – givenness marking, or a structural constraint that subsumes both inferable and repeated material – to determine which constituents must be deaccented. because of this unified constraint, it is not clear how the grammatical model would make such radically different predictions for the inferable conditions versus the repeated conditions. unlike the grammatical model, the accommodation model uses two distinct mechanisms to generate deaccenting on inferable and repeated material. deaccentuation of repeated material is derived grammatically by virtue of string identity with an antecedent. deaccentuation of inferable material is generated via an extragrammatical process of accommodation of an alternative antecedent containing an identical correlate to the deaccented material. the fact that there are separate mechanisms for inferable and repeated material means that the differential pattern of results proceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 181 https://doi.org/10.3765/elm https://www.elm-conference.net/ in these two conditions is not fundamentally incompatible with the accommodation model. experiments 2 and 3 tested the degree to which listeners’ preferences for accentuation or deaccentuation on inferable verbs were sensitive to manipulations in the broader discourse context. in experiment 2, the addition of a broader context supporting a “situationally identical” reading for inferable verbs and their antecedents decreased the magnitude of the preference for accentuation on inferable verbs. in experiment 3, addition of the presupposition trigger too reversed this preference to a preference for deaccentuation on inferable verbs. since the inferable verbs were already rated in the norming study as highly available even in the out-of-the-blue contexts, it is not clear how the grammatical model would explain the effect of these contextual manipulations on accent preferences. by contrast, the accommodation model allows for more holistic assessment of the discourse at large, with listeners potentially considering the entire context to determine whether the required accommodation is reasonable. deaccenting of inferable verbs was not rated as acceptable in out-of-the-blue contexts, but this pattern of results is not categorically ruled out by the accommodation model. further, the pragmatic mechanisms by which a supportive discourse context or a presupposition trigger could promote the necessary accommodation operation are relatively clear. 6. conclusion. in three experiments, we investigated listeners’ preferences for accentuation or deaccentuation on discourse-accessible or “inferable” verbs, and compared these preferences to the better-understood preferences for discourse-new and repeated constituents. contra commonly cited examples, deaccentuation of inferable constituents was rated as less natural than accentuation in out-of-the-blue contexts. this preference was mitigated somewhat with the addition of a discourse context supporting parallel readings of inferable verbs and their antecedents, and reversed to a preference for deaccentuation with the addition of the presupposition trigger too. we argued that the dual-mechanism accommodation model of deaccenting licensing is more compatible with the results of all three experiments than the single-mechanism grammatical account. future research exploring other aspects of this problem, such as the time course according to which accommodated inferences develop, may further confirm this conclusion. references baumann, stefan & arndt riester. 2012. referential and lexical givenness: semantic, prosodic and cognitive aspects. in gorka elordieta & pilar prieto (eds.), prosody and meaning, 1–34. berlin: mouton de gruyter. https://doi.org/10.1515/9783110261790.119. beaver, david i. & brady z. clark. 2008. sense and sensitivity: how focus determines meaning. malden, ma: wiley-blackwell. https://doi.org/10.1002/9781444304176. büring, daniel. 2007. semantics, intonation and information structure. in gillian ramchand & charles reiss (eds.), the oxford handbook of linguistic interfaces, 445–474. oxford: oxford university press. https://doi.org/10.1093/oxfordhb/9780199247455.013.0015. büring, daniel. 2016. intonation and meaning. oxford: oxford university press. https://doi.org/10.1093/acprof:oso/9780199226269.001.0001. chafe, wallace l. 1974. language and consciousness. language 50(1). 111–133. https://doi.org/10.2307/412014. chodroff, eleanor & jennifer cole. 2019. the phonological and phonetic encoding of information proceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 182 https://doi.org/10.3765/elm https://www.elm-conference.net/ structure in american english nuclear accents. in sasha calhoun, paola escudero, marija tabain & paul warren (eds.), proceedings of 19th international congress of phonetic sciences, melbourne, australia 2019, 1570–1574. canberra, australia: australasian speech science and technology association inc. chomsky, noam & morris halle. 1968. the sound pattern of english. new york: harper and row. van deemter, kees. 1994. what’s new? a semantic perspective on sentence accent. journal of semantics 11. 1–32. https://doi.org/10.1093/jos/11.1-2.1. van deemter, kees. 1999. contrastive stress, contrariety, and focus. in peter bosch & rob van der sandt (eds.), focus: linguistic, cognitive, and computational perspectives, 3–17. cambridge: cambridge university press. drummond, alex. 2020. ibex 0.3.7 manual. https://github.com/addrummond/ibex/ blob/master/docs/manual.md. fox, danny. 2000. economy and semantic interpretation. cambridge, ma: mit press. geiger, jeffrey & ming xiang. 2019. production of deaccenting under entailment, repetition, and bridging: phonetic and perceptual comparisons. in sasha calhoun, paola escudero, marija tabain & paul warren (eds.), proceedings of 19th international congress of phonetic sciences, melbourne, australia 2019, 512–516. canberra, australia: australasian speech science and technology association inc. lakoff, george. 1968. pronouns and reference. in james d. mccawley (ed.), syntax and semantics 7: notes from the linguistic underground, 275–335. new york: academic press. https://doi.org/10.1163/9789004368859 018. pierrehumbert, janet b. & julia hirschberg. 1990. the meaning of intonational contours in the interpretation of discourse. in philip r. cohen, jerry morgan & martha e. pollack (eds.), intentions in communication, cambridge, ma: mit press. https://doi.org/10.7551/mitpress/3839.003.0016. rochemont, michael s. 1986. focus in generative grammar. amsterdam: john benjamins. https://doi.org/10.1075/sigla.4. rooth, mats. 1992. ellipsis redundancy and reduction redundancy. in steve berman & arild hestvik (eds.), proceedings of the stuttgart ellipsis workshop: arbeitspapiere des sonderforschungsbereichs 340, stuttgart: universitäten stuttgart und tübingen in kooperation mit der ibm deutschland. schwarzschild, roger. 1999. givenness, avoidf and other constraints on the placement of accent. natural language semantics 7(2). 141–177. https://doi.org/10.1023/a:1008370902407. selkirk, elisabeth o. 1984. phonology and syntax: the relation between sound and structure. cambridge, ma: mit press. tancredi, christopher. 1992. deletion, deaccenting, and presupposition. cambridge, ma: mit dissertation. wagner, michael. 2012. focus and givenness: a unified approach. in ivona kučerová & ad neeleman (eds.), contrasts and positions in information structure, 102–147. cambridge: cambridge university press. https://doi.org/10.1017/cbo9780511740084.007. proceedings of elm 1: 172-183, 2021 jeffrey geiger and ming xiang: toward an accommodation account of deaccenting under nonidentity. 183 https://doi.org/10.3765/elm https://www.elm-conference.net/ intention and attention in image-text presentations: a coherence approach ilana torres, kathryn slusarczyk, malihe alikhani & matthew stone* abstract. in image-text presentations from online discourse, pronouns can refer to entities depicted in images, even if these entities are not otherwise referred to in a text caption. while visual salience may be enough to allow a writer to use a pronoun to refer to a prominent entity in the image, coherence theory suggests that pronoun use is more restricted. specifically, language users may need an appropriate coherence relation between text and imagery to license and resolve pronouns. to explore this hypothesis and better understand the relationship between image context and text interpretation, we annotated an image-text data set with coherence relations and pronoun information. we find that pronoun use reflects a complex interaction between the content of the pronoun, the grammar of the text, and the relation of text and image. keywords. elm; nlp; discourse; coherence; pronoun resolution; computational linguistics; semantics; pragmatics 1. introduction. image-text presentations are widely available on the internet, in captioned images, social media posts, and web pages. these image-text presentations provide a valuable proxy for situated language, enabling indirect inferences about face-to-face conversation, the primary setting for language learning and language use. mccullogh (2019) surveys the linguistic significance of using online communication to study spontaneous, informal language use. text and imagery function together in diverse ways (marsh & domas white 2003). an image of a dog posted on facebook relates to the caption, “this is my new puppy” in a way that is very unlike how an image of a model in a magazine relates to its caption “a model on a runway”. one fundamental difference is the semantic relationship between text and imagery: the model caption summarizes the image while the puppy caption links the image content to further facts about the speaker. these various relations lead to different ways in which we can identify objects in imagery through the use of a caption. a key case concerns the use of pronouns, which, in image-text presentations such as in the puppy image-caption example above, can refer deictically to entities from the image. pronouns occur often in text and conversation; they make utterances simpler and easier to process by eliminating the need to repeat a name or other descriptive content (see e.g., gordon and hendrick 1998). the semantic content of pronouns contains features such as number, gender, and person which helps in clarifying who or what a pronoun is referring to (büring 2011). however, extra-linguistic information such as real-life pointing can also be used to disambiguate a pronoun. when it comes to pronouns that are used in discourse, there is a further kind of information at hand that can be processed in order to resolve the pronoun: coherence relations (hobbs 1979). in particular, stojnic et al. (2013) argue that ambiguity of a pronoun in a text-image presentation can be resolved using coherence, by establishing specific inferential connections from the text to * this research was partly supported by nsf iis-1526723 and ccf-19349243. we thank the elm reviewers and attendees for comments and discussion that have improved the paper. authors: ilana torres, hofstra university (itorres2@hofstra.edu), kathryn slusarczyk, rutgers university (kat.slu@rutgers.edu), malihe alikhani, university of pittsburgh (malihe@pitt.edu), & matthew stone, rutgers university (matthew.stone@rutgers.edu). proceedings of elm 1: 273-283, 2021 c©2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone published by the lsa with permission of the author(s) under a cc by license. 273 https://doi.org/10.3765/elm https://www.elm-conference.net/ accompanying visual information that gives the reader or listener the context needed to identify the referent. while stojnic et al. (2013) examine video and accompanying narration, our work focuses on image-text pairs to allow for a closer analysis of the relationships between coherence relations and pronoun usage. this would mean that by processing discourse relations as we read a caption and regard the accompanying image, we are making use of relevant and important information which aids in resolving the (sometimes highly underspecified) content that can be found in captions. we can identify the referents of a pronoun by not only reading the caption but also by acknowledging what’s in the image. in previous work (alikhani et al 2019, alikhani et al 2020), we analyzed corpora of imagetext presentations to characterize their context-dependence as well as speakers’ communicative goals. in particular, for the annotation of image-text pairs in the conceptual captions data set of sharma et al (2018), we established a protocol to select types of coherence relations. the set of coherence relations we used included: (1) visible, (2) subjective, (3) action, (4) story, (5) meta, and (6) identification. examples of these relations from this dataset can be found in figure 1. further description of these relations from the current dataset can be found below under section 3.1., coherence relations. we used these coherence relations to capture how text applies to or relies on the accompanying image for information about context. this also allowed us to analyze these relations in terms of speakers’ communicative goals; the type of coherence relation and context provided is influenced by, and can indicate, what kind of information speakers intend to convey. figure 1: images and captions from a previous conceptual caption dataset as an example of initial coherence relations. (photo credits: yauhenka; danilo hegg) our previous work focused on coherence relations. here we expand the focus to consider pronouns. this has required a change of data set, not only to make sure that images feature salient objects, animals or people, but also to make sure that captions contribute appropriate coherence relations. previous annotations on discourse coherence relations in image-captioning have caused us to notice that there are higher correlations of pronouns occuring in story and subjective relations than in other relations. this was because speakers who use the story or subjective relations to describe their opinion about an image seem much more likely to draw on the prominence of entities in an visible, action, subjective action, story, meta caption: young happy boy swimming in the lake caption: approaching our campsite, at 1550m of elevation on the slopes proceedings of elm 1: 273-283, 2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone: intention and attention in image-text presentations: a coherence approach. 274 https://doi.org/10.3765/elm https://www.elm-conference.net/ image when formulating their utterance. in our current research on the usage of pronouns in imagetext pairs, we aim to examine how the types and frequency of pronouns used in captions is influenced by a caption’s coherence relation, and what this indicates about speaker intentions. we hypothesize that there is some pattern of correlation between image-caption discourse coherence relations and the types and frequency of pronouns within these captions. while we expect the highest frequency of all pronoun types to be in story and subjective type image-caption pairs, subjective type pairs in particular may show a higher frequency of using indexical pronouns like i, whereas in story relations we expect to see more examples of anaphoric pronouns. if any particular type of pronoun appears more often within certain types of coherence relations, or even in certain types of caption and utterance structures, we can draw links between image-captions, pronouns, and their references; these links may then offer insight into how speakers’ intentions affect pronoun usage, and vice versa. 2. methods. we created an interface to annotate a sample of image-text pairs. for each pronoun in the caption text, annotations were given on (1) discourse relation, (2) caption structure, and (3) pronoun type. we randomly sampled 6407 image-text pairs from the reddit dataset that all include pronouns. before beginning annotations, the first and second authors went through two rounds of preliminary annotations to adjust and finalize the annotation interface and establish strong interrater agreement. the first inter-rater agreement test we ran consisted of a set of 50 image-text pairs, with one or more pronouns per caption. this first test resulted in a low level of agreement, partially due to the inefficient first version of our caption structure types. we adjusted caption structure types to instead indicate utterance types and clarified pronoun distinctions between inter-raters. we reached a strong level of agreement with a second inter-rater agreement task and were able to continue with annotations. 3. annotation process. the annotators were presented with an image and the accompanying text along with options for choosing coherence relations, utterance structure, and pronoun type. 3.1. coherence relations. in our previous work on image-text coherence relations, we had modified existing coherence relations in order to fit the relationships we saw in our annotations. these relations were based on theoretical work on discourse coherence and structure (hobbs 1985, roberts 2012, webber 1999) as well as previous discourse annotation studies by prasad et al. (2008) and previous work by alikhani et al. (2019). as in our previous work, for each image we annotated, we chose one or more coherence relations based on the content of the text and its relation to the image. as listed above, the coherence relations were: (1) visible, where the content of the caption was depicted in the image, (2) subjective, where the caption was making a subjective statement about the content of the image, (3) action, where the caption describes a dynamic process of an action seen in the image, (4) story, where the caption provides a description of the image, or narrative-like background information, (5) meta, where the caption not only describes the image but also mentions productions and presentation of the image, and (6) identification where the caption uses a pronoun in order to identify a specific, salient object in the image. as mentioned, these relations are based on those previously used in text discourse; where visible relations are based on restatement relations, subjective relations on evaluation relations, action relations on elaboration relations, story relations on occasion relations, and meta relations on meta-talk relations (hobbs 1985, prasad et al. 2008). the identification relation was not present in our annotation guidelines for some previous work, as conceptual captions often have content omitted for machine learning experimentation. it was added in the current work given our specific inquiry into the usage of pronouns in image-text pairs. there was also an option for (7) irrelevant, proceedings of elm 1: 273-283, 2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone: intention and attention in image-text presentations: a coherence approach. 275 https://doi.org/10.3765/elm https://www.elm-conference.net/ which included images where the caption was gibberish or simply did not match the image, and (8) other, to indicate circumstances such as images which included text. an example of an irrelevant image-caption can be found in figure 2. further examples of coherence relations from the specific dataset can be found in figure 3. figure 2: example of an irrelevant image-caption. (photo credits: andre seale) figure 3: examples of various coherence relations. (photo credits: detap_rettiwt; ilana torres; alena capil) 3.2. utterance structures. the utterance structure type was also annotated to investigate the relationship between the structure of a caption and the frequency and types of pronouns within certain utterance structure types. with our first version of annotations for the structure of each caption, we agreed upon the following structure types; sentence, which indicated a full sentence regardless of punctuation; noun phrase with an implicit topic, with sub-categories for indicating whether the implicit topic was the image itself, the central focus of the image, or something else; and something else, to indicate a different structure. however, these types did not allow for caption: young girl walking on the dry grass field under daylight. irrelevant caption: he's not a purebred and he's not a puppy, but he's been my best friend for 12 years caption: my puppy smelling the flowers caption: it’s the most wonderful time of the year story, identification visible, action, identification subjective, story, meta proceedings of elm 1: 273-283, 2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone: intention and attention in image-text presentations: a coherence approach. 276 https://doi.org/10.3765/elm https://www.elm-conference.net/ meaningful annotation of captions that were not full sentences or noun phrases, as many captions included non-finite predicates. though an image of a kitten playing with a toy could be accompanied by the caption “my kitten is playing with her toy,” the shorter caption “playing with her toy” may also be used. annotation options were accordingly adjusted to include a wider range of structure types that appeared frequently in the dataset: (1) simple noun phrase, (2) noun phrase + non-finite predicate, (3) non-finite predicate, (4) full sentence, and (5) other, reserved for utterances like “ouch” that did not fall into the preceding annotation types. the first version of this annotation system allowed submission of just one annotation for each caption, but this made it difficult to accurately capture the structure of captions that appeared to contain multiple utterances, such as captions that contained both a full sentence and a predicate. we adapted our data collection to indicate the structure of each part of a caption, or each utterance, as we have designated them. while some captions were still treated as one utterance, those with punctuation that clearly defined separate sentences, phrases, or predicates were treated as multiple utterances. for example, a caption such as “this is my new puppy” would be treated as one utterance, while a caption such as “this is my new puppy. her name is lucky.” would be treated as two utterances, though the number of utterances within each caption was not noted. for each pronoun, we also annotated the structure of the utterance in which it appeared. 3.3. pronoun type. based on the definitions of pronouns in büring (2011) and traxler (2011), and the frequency of pronouns identified in previous analysis of coherence relations, we agreed upon the following categories for identifying pronoun type. we submitted an annotation for each pronoun in a caption. the options we agreed upon for pronoun annotations were (1) indexical (such as i and you), (2) demonstrative (such as this or that), (3) anaphoric (such as personal pronouns), (4) bound (such as bound personal pronouns), (5) indexical/bound (such as my and your), (6) backwards anaphora (such as a backward bound personal pronoun), and (7) not actually a pronoun, included to remove items that were mistakenly labeled as pronouns by the interface. 3.4. annotation process outline. we will use figure 4, below, as an example for a detailed outline of the annotation process. figure 4: (photo credits: annalise burke). caption: he's huge and lazy but when treats are involved, this big guy'll do anything proceedings of elm 1: 273-283, 2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone: intention and attention in image-text presentations: a coherence approach. 277 https://doi.org/10.3765/elm https://www.elm-conference.net/ • first, we identify the discourse coherence relations: story, subjective, and identification • next, we identify the caption structure: one full sentence; though this example includes punctuation, this is not a necessary condition of a full sentence annotation • lastly, we identify the pronouns: he is backwards anaphoric to this big guy, and this is demonstrative 4. results. overall, our dataset includes 13858 image-text pairs annotated with coherence relations out of which 6407 have pronouns. though this research is still in progress, our second inter-rater agreement task showed evidence that many of the sampled image-text pairs with pronouns fall into coherence relations of visible, meta, and story, as was evidenced in previous work. surprisingly, there were low levels of subjective captions. the overall distribution of coherence relations in the dataset can be seen in table 1. additionally, the most frequent pronouns overall were indexical and indexical/bound pronouns, followed by anaphoric. given that most captions were visible, meta, and story, the pronouns such as i, you, and other personal pronouns appeared very frequently. the distribution of pronouns in each coherence relation can be seen in table 2. the meta relation was particularly interesting, as other pronouns such as demonstrative pronouns were often found in captions with this relation. the distribution of pronouns in fine-grained meta captions can be found in table 3. though the distributions of each pronoun type appear to be similar across the fine-grained meta relation types, demonstrative pronouns appeared less frequently in meta-when relations than in meta-where and meta-how relations, and bound pronouns appeared more frequently in meta-how relations than in meta-where and meta-when relations. other findings include that, though not frequent, most cases of backwards anaphora appear in full sentence-annotated captions. table 4 shows the distribution of pronoun types in specific sentence structure types. additionally, table 5 indicates the distribution of sentence structures in captions containing specific coherence relations. our findings are discussed further below. visible subjective action meta story identification other 4014 (62.7%) 483 (7.53%) 936 (14.6%) 1998 (31.2%) 1463 (22.8%) 391 (6.1%) 785 (12.2%) table 1: the distribution of coherence relations in our dataset. the distribution of coherence relations for fine-grained meta categories of when, how and where are respectively 24.1%, 31.1%, and 63.3%. note that multiple coherence relations may be present in one example which explains why the sum of this row is greater than 100%. proceedings of elm 1: 273-283, 2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone: intention and attention in image-text presentations: a coherence approach. 278 https://doi.org/10.3765/elm https://www.elm-conference.net/ visible subjective action meta story identification indexical 29.92% 36.96% 31.75% 30.74% 34.14% 28.94% demonstrative 6.63% 8.26% 5.84% 10.12% 8.58% 9.13% anaphoric 13.72% 14.78% 13.50% 13.49% 13.47% 15.23% bound 8.06% 6.09% 7.66% 7.00% 5.82% 6.67% indexical_bound 32.98% 25.65% 35.40% 27.11% 28.10% 31.28% backanaphora 0.89% 1.30% 0.00% 1.04% 1.04% 1.36% other 0.04% 0.00% 0.00% 0.00% 0.05% 0.06% table 2: the distribution of pronouns in each category. each percentage indicates the texts containing pronouns of the indicated type as a percentage of the texts labeled with the indicated coherence relation. for example, 29.92% of image-text pairs annotated as visible contained at least one indexical pronoun. where when how indexical 30.58% 30.28% 25.00% demonstrative 13.28% 7.34% 12.50% anaphoric 12.78% 13.99% 12.50% bound 7.02% 7.57% 25.00% indexical_bound 25.81% 28.21% 25.00% backanaphora 0.75% 1.15% 0.00% other 0.00% 0.00% 0.00% table 3: the distribution of pronouns in fine-grained meta categories. as above, each percentage indicates the texts containing pronouns of the indicated type as a percentage of the texts labeled with the indicated fine-grained meta category. the notpronoun type indicates items that were incorrectly marked as pronouns by our annotation interface and will be disregarded in the following discussion. proceedings of elm 1: 273-283, 2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone: intention and attention in image-text presentations: a coherence approach. 279 https://doi.org/10.3765/elm https://www.elm-conference.net/ indexical demonstrative anaphoric bound back anaphora indexical bound np 8.51% 11.30% 13.30% 11.8% 12.00% 11.80% full sentence 80.30% 76.40% 77.40% 75.7% 84.00% 77.60% npnf predicate 10.40% 10.48% 8.71% 11.8% 4.00% 9.70% nf predicate 0.20% 01.31% 0.20% 0.60% 0.00% 0.10% other 0.40% 0.40% 0.30% 0.00% 0.00% 0.60% table 4: the distribution of pronoun types in sentence structure types. each figure indicates the utterance type containing the indicated coherence relation type as a percentage of all utterances containing the indicated coherence relation type. visible subjective action meta story np 11.5% 7.4% 8.6% 12.7% 9.8% full sentence 77.5% 76.2% 81.8% 78.1% 77.9% npnf predicate 10.2% 13.3% 9.1% 8.5% 11.0% nf predicate 0.3% 1.5% 0.0% 0.4% 0.4% other 0.3% 1.4% 0.2% 0.0% 0.6% table 5: the distribution of sentence structure types in coherence relation types. sentence structure type distribution for the identification relation is not listed as no images with an identification coherence relation have been annotated with sentence structure type yet. sentence structure types were introduced part way into the annotation process, and identification coherence relations are not very frequent, at only 9.9% of our annotated image caption pairs so far. proceedings of elm 1: 273-283, 2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone: intention and attention in image-text presentations: a coherence approach. 280 https://doi.org/10.3765/elm https://www.elm-conference.net/ 5. discussion. as we continue, our hypothesis still stands; that there is some pattern of correlation between image-caption discourse coherence relations and the types and frequency of pronouns within these captions. more than the overall distribution of coherence relation types in table 1, we are interested in the interactions of coherence relations, pronoun types, and sentence structures represented in tables 2 through 5. table 2 indicates that the most frequent types of pronouns overall are indexical, indexical/bound, and anaphoric pronouns; while each type seems to be about evenly represented across coherence relations, some less frequent and more frequent pairings are discussed below. as mentioned, our current results confirm that many of the sampled image-text pairs with pronouns fall into coherence relations of visible and story. given that pronouns in captions often refer to entities within the image, it is not surprising that visible is our most frequent relation at 62.7% of the annotated data set. of the data annotated as visible, the most frequent pronoun types were indexical/bound at 32.98% and indexicals at 29.92%. when a caption refers to entities like “my dog,” for example, “my” will require an indexical/bound annotation and “dog,” as long as a dog is pictured, will require a visible annotation. the frequency of these annotations is expected, since the current data set is composed of user generated images and captions that aim to describe the bound indexical relationship of the image’s main entity from the user’s perspective. as for story relations, the usage of any pronouns often give captions some element of backgrounded information that indicate their story relation. of the pronouns present in story relations, indexicals were the more frequent at 34.14%, with indexical/bound pronouns slightly behind at 28.10%. note that the most frequent and second most frequent pronoun types for visible relations and story relations are flipped, where images with visible relations are most often annotated with indexical/bound pronouns and then plain indexical pronouns, and images with story relations are most often annotated with plain indexical pronouns and then indexical/bound pronouns. indexical/bound pronouns like “my” (when used to reference a user’s dog, for example) can be taken as visible given the image of a dog, assuming that the dog must belong to someone and “my” is not necessarily an indicator of a story relation. indexicals like “i” or “you,” however, seemed to more often refer to entities that were not visibly within the image and therefore provided some information that cannot be verified for a visible annotation. this may explain why visible image-text pairs were slightly more often annotated with indexical/bound pronouns while story image-text pairs were slightly more often annotated with indexical pronouns. image-text pairs with demonstrative pronouns yielded some unexpected percentages. though we annotated demonstrative pronouns at similar rates (between 6.6% and 9.1%) for most coherence relation types, those with action coherence relations and meta (of any fine-grained type) coherence relations appeared at slightly differing frequencies of 5.84% and 10.12%, respectively. the lower frequency of demonstratives in action relations may be due to the preferred usage of indexical, indexical/bound, and anaphoric pronouns to refer to the entity taking action in the image. as for the higher rate of demonstratives in meta relations, we refer to the distributions in our fine-grained meta types in table 3, where demonstrative pronouns appeared less frequently in meta-when relations (7.34%) than in meta-where (13.28%) and meta-how (12.50%) relations. these higher frequencies seem to be indicative of how demonstrative pronouns like ‘this’ and ‘that’ can be used to refer to a place or some aspect of how an image was created, such as in ‘this photo’ or ‘that building.’ also dealing with the figures in table 3, bound pronouns appeared more frequently in meta-how relations than in meta-where and meta-when relations. captions with meta-how relations often appear to be more complex, disproportionately involving further clauses with coreference. proceedings of elm 1: 273-283, 2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone: intention and attention in image-text presentations: a coherence approach. 281 https://doi.org/10.3765/elm https://www.elm-conference.net/ additionally, the data set included subjective image-text pairs at a much lower frequency than we initially expected, at only about 7.53% of our data set. given the user generated source of the data set, we expected a higher frequency of subjective posts. however, the data set seemed to contain more objective visible captions, or those that simply stated other background information or related story captions. within the subjective image-text pairs we did have, the most frequent pronoun types were indexical pronouns at 36.96% and indexical/bound pronouns at 25.65%. though the third most frequent pronoun type is anaphoric at 14.78%, the remaining pronoun types were all below 9% of the total subjective image-text pairs annotated. this appears to be largely consistent with the other coherence relation types, though not all types have the same order of most frequent and second most frequent pronoun types. as mentioned, sentence structure types were introduced part way into the annotation process, meaning that the figures reported in tables 4 and 5 represent a smaller portion of the total data set. while full sentences were most frequent in image-text pairs using any given pronoun type, they were even more frequent in image-text pairs using backwards anaphora, at 84% of all images annotated with backwards anaphora. each of the sentence structure types have similar frequencies across the pronoun types, but backwards anaphora appeared relatively less frequently in noun phrases with non finite predicates, at only 4% of the category compared to an average of about 10% for other pronoun types. the high rate of backwards anaphoric pronouns in full sentences and lower rates in other sentence structures suggests that backwards anaphora is not efficient for captions with more truncated structures. table 5’s distribution of sentence structure types across coherence relations does not seem to show much besides a clear preference for full sentence type utterances; gleaning meaning from sentence structure type seems to require figures that include some information about pronoun types. additionally, we have not yet been able to report results for the distribution of sentence structure types in identification coherence relation image caption pairs yet. while the rate of identification relations in our full data set is low at 6.1%, the rate of subjective relations is similarly quite low at 7.53%. our dataset is biased as it doesn’t have balanced samples from each class of the relations or pronouns. we believe that a further expansion of the data set would allow us to report a distribution of sentence structure types within all coherence types. 6. conclusion. we found that pronoun use depends on the kind of relation between the image and its caption. we saw that there is overall a high frequency of visible coherence relations, and the most frequently, indexical and indexical/bound personal pronouns occurred in captions, followed by anaphoric pronouns. the kind of sentence structure type used in a caption also correlated with pronoun usage: for example, backward anaphora was most common in full sentences. our additional annotation of utterance structures may reveal further related patterns between coherence relations, structure, and pronoun use, and allow us to analyze how image-text pairs are created according to speaker intentions. our current research provides opportunities for future work on pronoun resolution in the context of image-captioning, and will allow the construction of more accurate and effective captioning models, which will assist in the creation of model-generated image captioning as well as better results for search engines. these models will ideally be able to create strong captions for given images, thanks to research on the content of captions and their visual referents. however thorough this research on coherence relations in the english language may be, this leaves room for research on image-text relations in various other languages. different languages have different paradigms for usage of pronouns or a complete lack of pronouns, and similar annotations such as from this experiment would allow a better understanding of how languages process pronouns in small segments of language and especially in addition to images. proceedings of elm 1: 273-283, 2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone: intention and attention in image-text presentations: a coherence approach. 282 https://doi.org/10.3765/elm https://www.elm-conference.net/ additionally, there might be cultural differences that arise in image-caption pairs posted in different languages which could be studied as well. our dataset is available on the project github page.1 references alikhani, m., chowdhury, s. n., de melo, g., & stone, m. (2019). a corpus of image-text discourse relations. proceedings of the 2019 conference of the north american chapter of the association for computational linguistics: human language technologies, 1, 570575. alikhani, m., sharma, p., li, s., soricut, r., stone, m. (2020). cross-modal coherence modeling for caption generation. proceedings of the 58th annual meeting of the association for computational linguistics. büring, d. (2011). pronouns. in von heusinger, maienborn and portner (eds.): semantics: an international handbook of natural language meaning, 2, 971-995. gordon, p. & hendrick, r. (1998). the representation and processing of coreference in discourse. cognitive science, 22(4), 389-424. hobbs, j. r. (1979). coherence and coreference. cognitive science, 3(1), 67-90. hobbs, j. r. (1985). on the coherence and structure of discourse. technical report, sri international menlo park ca. marsh, e. e. & domas white m. (2003). a taxonomy of relationships between images and text. journal of documentation, 59(6), 647-672. mcculloch, g. (2020). because internet. penguin usa. prasad, r., dinesh, n., lee, a., miltsakaki, e., robaldo, l., joshi, a. k., & webber, b. l. (2008). the penn discourse treebank 2.0. in lrec. citeseer roberts, c. (2012). information structure: towards and integrated formal theory of pragmatics. semantics and pragmatics, 5, 6-1. sharma, p., ding, n., goodman, s., & soricut, r. (2018). conceptual captions: a cleaned, hypernymed, image alt-text dataset for automatic image captioning. proceedings of the 56th annual meeting of the association for computational linguistics, 1, 2556-2565. shiffrin, d. (1980). meta-talk: organizational and evaluative brackets in discourse. sociological inquiry, 50(3-4), 199-236. stojnic, u., stone, m., & lepore, e. (2013). deixis (even without pointing). (report). philosophical perspectives, 27(1). traxler, m. j. (2011). introduction to psycholinguistics: understanding language science. wileyblackwell. webber, b., knott, a., stone, m., & joshi, a. (1999). discourse relations: a structural and presuppositional account using lexicalised tag. proceedings of the 37th annual meeting of the association for computational linguistics on computational linguistics, 41-48. 1 https://github.com/malihealikhani/elm2020-intention-and-attention-in-image-textpresentations proceedings of elm 1: 273-283, 2021 ilana torres, kathryn slusarczyk, malihe alikhani and matthew stone: intention and attention in image-text presentations: a coherence approach. 283 https://doi.org/10.3765/elm https://www.elm-conference.net/ visual boundaries in sign motion: processing with and without lip-reading cues julia krebs, evie a. malaia, ronnie b. wilbur & dietmar roehm* abstract. sign languages demonstrate a higher degree of iconicity than spoken languages. studies on a number of unrelated sign languages show that the event structure of verb signs is reflected in the phonological form of the signs (wilbur (2008), malaia & wilbur (2012), krebs et al. (2021)). previous research showed that hearing nonsigners (with no prior exposure to sign language) can use the iconicity inherent in the visual dynamics of a verb sign to correctly identify its event structure (telic vs. atelic). in two eeg experiments, hearing non-signers were presented with telic and atelic verb signs unfamiliar to them, which they had to classify in a two-choice lexical decision task in their native language. the first experiment assessed the timeline of neural processing mechanisms in non-signers processing telic/atelic signs without access to lip-reading cues in their native language, to understand the pathways for incorporation of physical perceptual motion features into linguistic processing. the second experiment further probed the impact of visual information provided by lip-reading (speech decoding based on visual information from the face of the speaker, most importantly, the lips) on the processing of telic/atelic signs in non-signers. keywords. semantics; eeg; psycholinguistics; sign language; event visibility 1. introduction. in the course of human evolution, the ability to identify and interpret discrete events in the fluidly changing environment was one of the most critical functions of cognition. as humans developed the ability to communicate using language, information about actions their structure, temporal parameters, and participants took a central role in linguistic communication in the form of verbs and their linguistic features. verbs and their arguments are central to any communicative message, and consistencies in their relationships provide the basis of linguistic patterns across languages (evans & levinson (2009), greenberg et al. (1963)). every sentence in linguistic communication is centered on transmitting information about an action or an event, that is, predication. the verb and its arguments, which provide the basis of every sentence, can describe the event in two ways: as having an inherent boundary or an endpoint (telic events), or not inherently bounded or limited (atelic events). understanding how action processing feeds into language processing can be groundbreaking in terms of modeling language disorders, identifying them early, and developing therapies. the hypothesis that language builds on general, non-linguistic abilities such as the ability to identify, parse, and interpret actions has not been conclusively tested in spoken languages, as they differ from action in modality (auditory vs. visual). sign languages allow investigation of the processes of action comprehension and language understanding within a single modality, testing the relationship between the two at various processing stages, from sensory perception to higher cognition. *authors: julia krebs, university of salzburg (julia.krebs@plus.ac.at), evie a. malaia, university of alabama (eamalaia@ua.edu), ronnie b. wilbur, purdue university (wilbur@purdue.edu) & dietmar roehm, university of salzburg (dietmar.roehm@plus.ac.at). proceedings of elm 2: 164-175, 2023 c©2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm published by the lsa with permission of the author(s) under a cc by license. 164 https://doi.org/10.3765/elm https://www.elm-conference.net/ in sign language verbs, event structure is often perceptually reflected in the form of the signs, i.e. the hand articulator motion dynamics. for example, wilbur (2003) observed that in american sign language (asl) lexical verbs can be analyzed as telic and atelic based on their phonological form, with telics having a more rapid deceleration to the place of articulation at the end of the sign reflecting semantic end-state of affected arguments. the observation that semantic verb classes are characterized by certain movement profiles was formulated as the event visibility hypothesis (evh; wilbur (2008)). empirical evidence for the evh came from motion capture research, which indicated systematic kinematic distinctions between telic and atelic verbs, whereby the endpoint of the event in telics is marked by a higher peak velocity and significantly faster deceleration at the end in contrast to atelics (malaia et al. (2008, 2013a), krebs et al. (2021), malaia & wilbur (2012)). prior research suggests that from the standpoint of neural computations, language and action processing have a lot of overlap. humans rely on dynamic features of visual motion for perceptual segmentation of the visual and linguistic signal. multiple studies have shown that reality is segmented into events at multiple scales simultaneously (zacks et al. (2001a,b)). such event segmentation studies typically ask participants to watch a video with a dynamic scene and indicate time-points at which the participants think an action is completed; participants can do so at finegrained and coarse-grained boundaries. across cohorts, participants show remarkable agreement in identifying the timing boundaries of both coarse and fine-grained events, either in realistic scenarios (e.g. how one folds laundry), or in abstract moving-dot experiments (kurby & zacks (2008), speer et al. (2007), zacks et al. (2001a)). the ability to identify, hierarchically structure, and remember segmented portions of the signal appears to be transferable between action and linguistic domains. strickland et al. (2015) provided an example of action-to-language processing transfer, showing that non-signers are capable of identifying telic/atelic semantics of sign language verbs in the absence of any prior exposure to a sign language. non-signers, who were shown videos of sign language verbs differing in event structure and resulting motion signatures, were asked to select the likely meaning of the observed sign from two english verbs. participants accurately inferred lexical aspectual meaning (’aktionsart’) from visual stimuli, distinguishing between atelic and telic signs with unknown meaning. the fact that non-signers were able to make sense of the visual signal suggests that the presence/absence of a dynamic visual boundary was sufficient for action segmentation. due to the linguistic nature of the task, inference about the event structure of the verb would have been carried out on the basis of action segmentation (telic vs. atelic). strickland et al. (2015) concluded that linguistic notions of telicity and mapping biases between telicity and visual form were universally accessible, shared between signers and non-signers. a reverse phenomenon language-to-action transfer of skill in segmentation of visual signal has been demonstrated in a series of experiments in which signers and non-signers were asked to reproduce dynamic point-light drawings (klima et al. (1999)). signers, but not speakers, made a crucial distinction between strokes and transitions in the point-light display: signers did not draw the lines which represented transitional motion between “strokes” of the drawings. the stimuli were not linguistically informative for any of the participants; however, the signing participants were able to extrapolate their linguistic experience in visual segmentation of a signal (e.g. ignoring proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 165 https://doi.org/10.3765/elm https://www.elm-conference.net/ transitional movements between meaningful signs) to a non-linguistic task that focused on action segmentation and structuring. while non-signers and signers appear capable of relying on similar motion cues for segmentation of the visual signal and assignment of meaning, signers are capable of more nuanced structuring of the visual signal. from multiple perceptual features experimentally tested as potentially relevant for visual action comprehension (e.g. distance between pairs of moving objects, relative location, speed, acceleration, etc.), changes in speed of individual objects emerged as the feature most highly correlated with event boundary identification. action start and end times, as identified by participants, are highly correlated with increases and decreases of speed (acceleration and deceleration) (zacks et al. (2006)). rate of deceleration is also one of the motion features used for differentiating telic from atelic verbs in sign language production. at the neural level, these changes in speed of individual objects were associated with increased activity in the area of the brain termed mt+, and a nearby region in the superior temporal sulcus – both associated with processing of biological motion (zacks et al. (2006)). very similar neural activations were reported in sign-naı̈ve participants observing signed sentences in asl involving telic and atelic verbs (malaia et al. (2012a)); yet, signers observing the same stimuli show focused activation in the left inferior frontal gyrus, an area related specifically to language processing. this indicates that while both signers and non-signers operate on the same perceptual information (i.e. both visually process the perceptual-kinematic difference between telic and atelic asl signs), only familiarity with the language allows low-level perception of motion differences in the signal to be processed as information at the linguistic levels. the findings described so far show that a perceptual-kinematic velocity feature used for nonlinguistic event segmentation is incorporated into the language system to be processed as an abstract linguistic feature by deaf 1 signers (malaia et al. (2012a)). although the cross-linguistic nature of motion-based interpretation of lexical aspect in sign languages is widely attested (wilbur (2008)), along with the consistency of both signers and non-signers in interpreting signs with such features (strickland et al. (2015), kuhn et al. (2021)), the neural bases of this universal mapping from motion features to linguistic features are not well-described. to investigate the neural timeline of mapping between visual motion and linguistic event structure, we recorded erp (event related potential) data during processing of telic and atelic signs in hearing non-signers. participants were asked to label the viewed signs using a two-alternative-forced-choice task in their native language, and, additionally, to indicate how certain they were of their decision. in experiment 1, the sign language stimuli, which represented unknown input for the participants, consisted of signed telic and atelic verbs from turkish sign language (tid), italian sign language (lis), sign language of the netherlands (ngt) (from strickland et al. (2015)), and croatian sign language (hzj). in experiment 2, the sign language stimuli consisted of telic and atelic signs from austrian sign language (ögs) which were accompanied by mouthing (mouth movement forming (part of) a german word) that potentially provided additional information to the participants who were native german speakers. based on previous research (strickland et al. (2015), kuhn et al. (2021)), we hypothesized that non-signers would be able to accurately classify telic/atelic verbs. in line with ji & papafragou 1per convention deaf with upper-case d refers to deaf or hard of hearing humans who define themselves as members of the sign language community. in contrast, deaf refers to audiological status. proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 166 https://doi.org/10.3765/elm https://www.elm-conference.net/ (2020), we expected to see higher classification accuracy for bounded events. previous neurolinguistic studies also showed that hearing non-signers relied on sensory/occipital cortices (including mt+ region) when processing telic vs. atelic signs (malaia et al. (2012b)). we thus expected that the sensory-perceptual difference between verb types would be reflected on the neurophysiological level in erps within early time windows (before 300 msec post-stimulus onset). our research question centered on the processing mechanisms involved in the form-to-meaning mapping/integration process (past 300 msec post-stimulus onset). due to the linguistic nature of the task, linguistic processing indicators could be expected in both conditions. however, based on prior behavioral research, it could be expected that the timeline for the process of linguistic mapping/integration might differ between telic and atelic signs within each experiment, as well as between the experiments with and without lip-reading cues. 2. methods. 2.1. participants. 27 participants (21 female) were included in the final analysis, with a mean age of 22.96 years (sd = 3.98; range = 16-31 years). all of them were hearing students without competence in any sign language and all of them were right-handed (tested by an adapted german version of the edinburgh handedness inventory; oldfield (1971)). at the time of the study none showed any neurological or psychological disorders. all had normal or corrected vision and were not influenced by medication or other substances which may impact cognitive ability. the participants either received 20c or got credits for their study program. 2.2. materials and design. experiment 1: a 1 x 2 design with the two-level factor telicity involving telic and atelic signs was used. 36 verbs were presented in each condition (72 critical verbs), with 92 fillers, resulting in a total of 144 items. the stimuli consisted of signs that were used in the study of strickland et al. (2015), that is atelic and telic verbs from tid, lis and ngt, supplemented by signs from hzj to achieve the appropriate stimuli number. for each of the sign languages, 9 atelic and 9 telic verbs were presented2. experiment 2: equivalent to experiment 1, but the stimuli consisted of 36 atelic and 36 telic signs of ögs3. 2.3. procedure. the material was presented in 6 blocks (24 verbs in each block). every trial started with the presentation of a stimulus video presented in the middle of the screen with a size of 820 x 540 (25 fps). the video was followed by a two-choice decision task, similar to the labeling task used by strickland et al. (2015). participants were asked to guess the meaning of the presented sign by forced-choice selection from two answer choices in written german. to ensure that the participants could not determine the meaning of the signs by iconically relating the meaning to the sign in experiment 1, both answers did not show the meaning of the presented sign, but one answer choice matched the stimulus with respect to telicity. in experiment 2, one of the answers matched both the semantics and event structure (telicity) of the stimuli, while the other had the opposite event structure. one telic and one atelic verb were presented. after the labeling task the 2the mean length of the critical stimulus videos as well as standard deviation and range of video length per condition (values given in seconds): atelics: mean: 2.06; sd: 0.47; range: 1-3; telics: mean: 1.75; sd: 0.55; range: 1-3. 3the mean length of the critical stimulus videos as well as standard deviation and range of video length per condition (values given in seconds): atelics: mean: 2.75; sd: 0.44; range: 2-3; telics: mean: 2.61; sd: 0.49; range: 2-3. proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 167 https://doi.org/10.3765/elm https://www.elm-conference.net/ participants rated how certain they were of their decision on a 7 point likert scale (one stands for “very unsure”, four means “about 50% sure” and seven indicates “very sure”). prior to the experiment, a training block was presented to familiarize participants with task requirements and permit them to ask questions. the duration of breaks after each block was determined by the participants themselves. participants were instructed to avoid eye movements and other motions during the presentation of the video material. the participants filled out a written questionnaire containing demographic questions and questions relevant for eeg data recording. informed consent was obtained in written form. the eeg was recorded from twenty-six electrodes (fz, cz, pz, oz, f3/4, f7/8, fc1/2, fc5/6, c3/4, cp1/2, cp5/6, p3/7, p4/8, o1/2, po9/10) fixed on the participant’s scalp by means of an elastic cap (easy cap, herrsching-breitbrunn, germany). horizontal eye movements (heog) were registered by electrodes at the lateral ocular muscles (left and right) and vertical eye movements (veog) were recorded by electrodes fixed above and below the left eye. all electrodes were referenced against the electrode on the left mastoid bone and offline re-referenced against the averaged electrodes at the left and right mastoid. the afz electrode functioned as the ground electrode. the eeg signal was recorded with a sampling rate of 500 hz. for amplifying the eeg signal we used a brain products amplifier (high pass: 0.01 hz). in addition, a notch filter of 50 hz was used. the electrode impedances were kept below 5 kω. offline, the signal was filtered with a bandpass filter (butterworth zero phase filters; high pass: 0.1 hz, 48 db/oct; low pass: 20 hz, 48 db/oct). 3. data analysis. 3.1. behavioral data. experiment 1: the effects of telicity and language were examined for the participants’ accuracy regarding the two-choice decision task. behavioral data per participant and per item were assessed using repeated-measures analysis of variance (anova). the fixed factors telicity (telic vs. atelic) and language (lis, tid, hzj, ngt) and the random factors subjects (fsubj) and items (fitem) were included. the statistical analysis was carried out hierarchically; only significant interactions (p≤.05) were resolved using a step-down approach. to correct for violations of sphericity, the greenhouse & geisser (1959) correction was applied to repeated measures with greater than one degree of freedom. only significant effects (p≤.05) are reported. experiment 2: analysis was the same as in experiment 1, except that only the fixed factor telicity (telic vs. atelic) was included in the analysis. 3.2. erp data. data analysis was the same for the two experiments. to determine the onset and offset of the effects, we computed a 50 msec time window analysis. statistical evaluation of the erp data was carried out by comparison of the mean amplitude of the erps within the time window, per condition and per subject in two regions of interest (rois). the factor roi involved the levels anterior = f3, f4, f7, f8, fc1, fc2, fc5, fc6, fz, cz, and posterior = p3, p4, p7, p8, po9, po10, o1, o2, pz, oz. the signal was corrected for ocular artifacts by the gratton and coles method (gratton et al. (1983)) and screened for artifacts (minimal/maximal amplitude at -75/+75 µv). data was baseline-corrected to -300 to 0. statistical analysis was carried out in a hierarchical manner, that is, only significant interactions (p≤.05) were included in a step-down analysis. for statistical analysis of the erp data an anova was computed including the factors proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 168 https://doi.org/10.3765/elm https://www.elm-conference.net/ of condition telicity (atelic vs. telic) and roi. only significant effects (p≤.05) are reported. erps were measured with respect to the time point when the target handshape reaches the target location where the movement of the verb sign starts. 4. results. 4.1. experiment 1: no lip-reading cues. 4.1.1. behavioral data. participants gave correct responses above chance level regarding telics and atelics and were more accurate regarding telics in most of the languages (telic: hzj: 73.25%, lis: 79.42%, ngt: 72.43%, tid: 64.20% accuracy; atelic: hzj: 53.09%, lis: 63.79%, ngt: 53.50%, tid: 73.66% accuracy). the anova of participants’ accuracy regarding the twochoice decision task revealed a significant main effect of telicity [fsubj(1,26)=37.13, p<.001, η2p=.59], a significant main effect of language [fsubj(3,78)=5.49, p<.002, η2p=.17] and a significant interaction telicity x language [fsubj(3,78)=12.04, p<.001, η2p=.32]. the resolution of the interaction by language revealed significant telicity effects for hzj [f(1,26)=23.57, p<.001, η2p=48], lis [f(1,26)=19.72, p<.001, η2p=.43], tid [f(1,26)=8.57, p<.007, η2p=.25] and ngt [f(1,26)=15.97, p<.001, η2p=.38]. 4.1.2. erp data. significant processing differences for telics compared to atelics were revealed at the neurophysiological level. beginning from sign onset (i.e. target handshape positioned in target location), statistically significant neural differences in processing appeared across several time ranges anteriorly (0-200 msec, 500-550 msec, 650-800 msec, 850-1300 msec, and 14001500 msec), posteriorly (600-700 msec, 750-1050 msec, and 1250-1300 msec), and in a broadly distributed manner (200-250 msec and 300-400 msec) (see figure 1). 4.2. experiment 2: lip-reading cues. 4.2.1. behavioral data. participants gave correct responses above chance level regarding telic and atelic signs. they were more accurate with respect to the telic condition compared to the atelic condition (telic, 94.75% accuracy; atelic, 89.81% accuracy). the analysis of variance of participants’ accuracy revealed a significant main effect of telicity [fsubj(1, 26) = 22.49, p <.001, η2p = .46]. 4.2.2. erp data. with erp onset time-locked to the point when the target handshape of the sign reached target location, data analysis revealed a more posteriorly distributed positive effect for telic compared to atelic signs in the 250 to 500 msec time window, a broadly distributed positive effect in the 500 to 600 msec time window, and a posteriorly distributed positive effect in the 600 to 1650 msec time window. furthermore, an anteriorly distributed negative effect for telic compared to atelic signs was identified in the 1800 to 1850 msec window (see figure 2). 5. discussion. replicating previous results, the behavioral data analysis indicates that non-signers classify signs as telic or atelic with high accuracy (strickland et al., 2015; kuhn et al., 2021). across sign languages, telic signs were classified more accurately than atelic signs. the finding that participants classified telics more accurately than atelics is in line with ji & papafragou (2020), who report that the category of bounded events was identified with greater ease compared to that of unbounded events. ji & papafragou (2020) suggest that the involvement of an internal structure that culminates in defined endpoints makes bounded events easier to individuate, track, and generalize proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 169 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: telic (red)/atelic (blue) sign processing without non-manual cues; difference wave in black over, as compared to unbounded events. the present data extends this observation to sign language stimuli and the use of a linguistic task. differences between processing of telic and atelic signs were also found at the neurophysiological level, since different erp patterns were observed for experiment 1 and experiment 2. in experiment 1, erp analysis identified differences in processing timeline between telic vs. atelic proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 170 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2: telic (red)/atelic (blue) sign processing with non-manual cues in native language; difference wave in black signs in both early (prior to 300 msec past stimulus onset), and later time windows. the effects in early time windows (starting at sign onset) likely reflect the difference in sensory-perceptual processing, i.e. the processing of the difference in movement dynamics between verb types. anterior and posterior erp effects for telic compared to atelic stimuli appearing in later time windows likely reflect different mapping/integration processes for telic signs. experiment 2 erp data indicated proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 171 https://doi.org/10.3765/elm https://www.elm-conference.net/ later onset of processing differences between telic vs. atelic signs, and almost exclusively posterior (temporo-parietal-occipital) distribution of sustained (to 1600 msec) differences in processing. in debriefing, the participants indicated their reliance on mouthing information for this experiment, which suggests likelihood of attempts at integrating visual speech (lip-reading), manual, and spoken lexical information prior to the decision-making task. in contrast to experiment 1, experiment 2 revealed erp effects for telic compared to atelic signs that started in later time windows, extended into later time windows, and showed a primarily posterior distribution. thus, instead of the early perceptual processing based on sign kinematics, observed in experiment 1, the participants seemed to rely on mouthing information, as described in the lip-reading literature. for example, research on lip-reading of speech in silence shows that observation of lip-motion leads to generation of auditory speech representation in temporal auditory cortices (bourguignon et al. (2020)). the sustained parietal effects in experiment 2, then, might be due to multimodal (visual and auditory) stream integration, as well as, possibly, lexical access resulting from successful integration; however, more specific research is necessary to ascertain the timeline of audiovisual integration for sign language mouthing in non-signers. the differences in the morphology of erp effects elicited by telic vs. atelic stimuli likely reflect differences in cognitive/linguistic processing between verb types and different mapping and/or integration processes for telics compared to atelics. previous work showed that the telicity in verbs may facilitate online language processing, for example, in resolution of garden path sentences. malaia et al. (2009, 2012b, 2013b) investigated the effects of verbal telicity on syntactic reanalysis of reduced relative clauses in written english, whereby the verb in the relative clause was either telic or atelic. sentences with atelic signs imposed higher processing costs at the disambiguation point, as compared to sentences with telic signs. reduced relative clauses required re-assignment of thematic roles, which appeared to proceed more rapidly in sentences with telic verbs, potentially because bounded verbs triggered extraction of event template along with thematic roles inherent in it, thereby facilitating thematic role re-assignment. atelic verbs, which did not provide the conceptual boundary for event segmentation, did not appear to trigger the same processing mechanism. in our experiments, participants viewed unfamiliar signs, followed by an offline classification task. erp effects thus might reflect the segmentation operation the participants carried out for visually bounded telic signs. although the segmentation operation might require more effort (e.g. attentional allocation and memory reference) at the point of being carried out, it is likely to facilitate the participants’ performance in the offline classification task later (malaia et al. (2009), ji & papafragou (2020)). therefore, the online erp effects for telic signs, as compared to atelic signs, might potentially stem from two different sources: recruitment of additional processing resources for telics in the segment toward the end of each sign, or release of cognitive resources past sign offset. crucially, the difference is observable at the neurophysiological level, suggesting that action-to-language mapping processes differ between visually bounded and unbounded events. the observed differences regarding behavioral and erp results observed for both studies suggest that participants used a different strategy in the two experiments. whereas in experiment 1, non-signers seemed to segment the visual signal on the basis of the signs’ motion profiles, experiment 2 suggests that, if available, non-signers use lip movement information visual cues proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 172 https://doi.org/10.3765/elm https://www.elm-conference.net/ which they are familiar with from their l1 for classifying unknown signs. thus, in experiment 2 non-signers paid more attention to lip-reading (as self-reported after the experiment), as opposed to tracking visual motion profiles in the stimuli. because linguistic information provided by lip movement is part of audio-visual spoken language processing, it was easier for non-signers to classify the signs in experiment 2 compared to experiment 1. these findings might reflect the potential evolutionary pathway of how physical-perceptual motion features were co-opted into the linguistic structure of sign languages. cross-linguistic similarities in the visual representation of event structure have been described for a number of unrelated sign languages (malaia & milković (2021), krebs et al. (2021)). sign languages, however, differ in linguistic representation of event structure, such that realization of end-state marking might take on various forms. comparative analysis of motion capture data also points to a variety of strategies for the mapping between physical parameters for articulator motion, and linguistic features that incorporate boundedness. thus, although sign languages mark event structure in an iconic way, they show language-specific characteristics with respect to how event structure is represented and expressed. however, despite these differences, non-signers can classify these iconically motivated forms accurately, because articulator motion profiles overall are similar to motion profiles of observed events. this finding provides further neurophysiological evidence for the event segmentation theory in perception (zacks & tversky (2001), zacks & swallow (2007)) and the evh for sign languages (wilbur (2003, 2008, 2010)). acknowledgements we want to thank brent strickland and his colleagues for providing their video material, marijke scheffener and roland pfau for signing and filming the ngt videos, brigita sedmak and marina milković for filming and providing the hzj videos, and waltraud unterasinger for signing the video material in austrian sign language. many thanks to all participants taking part in this study. this work was supported in part by the national science foundation (nsf) award #1734938 and the austrian science fund (fwf): 35671. references bourguignon, mathieu, martijn baart, efthymia c kapnoula & nicola molinaro. 2020. lipreading enables the brain to synthesize auditory features of unknown silent speech. journal of neuroscience 40(5). 1053–1065. evans, nicholas & stephen c levinson. 2009. the myth of language universals: language diversity and its importance for cognitive science. behavioral and brain sciences 32(5). 429–448. gratton, gabriele, michael gh coles & emanuel donchin. 1983. a new method for off-line removal of ocular artifact. electroencephalography and clinical neurophysiology 55(4). 468– 484. greenberg, joseph h et al. 1963. some universals of grammar with particular reference to the order of meaningful elements. universals of language 2. 73–113. greenhouse, samuel w & seymour geisser. 1959. on methods in the analysis of profile data. psychometrika 24(2). 95–112. ji, yue & anna papafragou. 2020. is there an end in sight? viewers’ sensitivity to abstract event structure. cognition 197. 104197. klima, edward s, ovid jl tzeng, yya fok, ursula bellugi, david corina & jeffrey g bettger. proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 173 https://doi.org/10.3765/elm https://www.elm-conference.net/ 1999. from sign to script: effects of linguistic experience on perceptual categorization. journal of chinese linguistics monograph series 96–129. krebs, julia, gerda strutzenberger, hermann schwameder, ronnie wilbur, evie malaia & dietmar roehm. 2021. event visibility in sign language motion: evidence from austrian sign language. in proceedings of the annual meeting of the cognitive science society, vol. 43, 362–368. kuhn, jeremy, carlo geraci, philippe schlenker & brent strickland. 2021. boundaries in space and time: iconic biases across modalities. cognition 210. 104596. kurby, christopher a & jeffrey m zacks. 2008. segmentation in the perception and memory of events. trends in cognitive sciences 12(2). 72–79. malaia, evguenia, john borneman & ronnie b wilbur. 2008. analysis of asl motion capture data towards identification of verb type. in semantics in text processing. step 2008 conference proceedings, 155–164. malaia, evguenia, ronnie b wilbur & christine weber-fox. 2009. erp evidence for telicity effects on syntactic processing in garden-path sentences. brain and language 108(3). 145– 158. malaia, evie & marina milković. 2021. aspect: theoretical and experimental perspectives. in j. quer, r. pfau & a. herrmann (eds.), the routledge handbook of theoretical and experimental sign language research, 194–212. routledge. malaia, evie, ruwan ranaweera, ronnie b wilbur & thomas m talavage. 2012a. event segmentation in a visual language: neural bases of processing american sign language predicates. neuroimage 59(4). 4094–4101. malaia, evie & ronnie b wilbur. 2012. kinematic signatures of telic and atelic events in asl predicates. language and speech 55(3). 407–421. malaia, evie, ronnie b wilbur & marina milković. 2013a. kinematic parameters of signed verbs. journal of speech, language, and hearing research 56(5). 1677–1688. malaia, evie, ronnie b wilbur & christine weber-fox. 2012b. effects of verbal event structure on online thematic role assignment. journal of psycholinguistic research 41(5). 323–345. malaia, evie, ronnie b wilbur & christine weber-fox. 2013b. event end-point primes the undergoer argument: neurobiological bases of event structure processing. in studies in the composition and decomposition of event predicates, 231–248. springer. oldfield, richard c. 1971. the assessment and analysis of handedness: the edinburgh inventory. neuropsychologia 9(1). 97–113. speer, nicole k, jeffrey m zacks & jeremy r reynolds. 2007. human brain activity time-locked to narrative event boundaries. psychological science 18(5). 449–455. strickland, brent, carlo geraci, emmanuel chemla, philippe schlenker, meltem kelepir & roland pfau. 2015. event representations constrain the structure of language: sign language as a window into universally accessible linguistic biases. proceedings of the national academy of sciences 112(19). 5968–5973. wilbur, ronnie. 2003. representations of telicity in asl. in proceedings from the annual meeting of the chicago linguistic society, vol. 39 1, 354–368. chicago linguistic society. wilbur, ronnie b. 2008. complex predicates involving events, time and aspect: is this why sign proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 174 https://doi.org/10.3765/elm https://www.elm-conference.net/ languages look so similar. in j. quer (ed.), signs of the time: selected papers from tislr 2004, 219–250. signum. wilbur, ronnie b. 2010. the semantics-phonology interface. in d. brentari (ed.), cambridge language surveys: sign languages, 357–382. cambridge: cambridge university press. zacks, jeffrey m, todd s braver, margaret a sheridan, david i donaldson, abraham z snyder, john m ollinger, randy l buckner & marcus e raichle. 2001a. human brain activity timelocked to perceptual event boundaries. nature neuroscience 4(6). 651–655. zacks, jeffrey m & khena m swallow. 2007. event segmentation. current directions in psychological science 16(2). 80–84. zacks, jeffrey m, khena m swallow, jean m vettel & mark p mcavoy. 2006. visual motion and the neural correlates of event perception. brain research 1076(1). 150–162. zacks, jeffrey m & barbara tversky. 2001. event structure in perception and conception. psychological bulletin 127(1). 3–21. zacks, jeffrey m, barbara tversky & gowri iyer. 2001b. perceiving, remembering, and communicating structure in events. journal of experimental psychology: general 130(1). 29–58. proceedings of elm 2: 164-175, 2023 julia krebs, evie a. malaia, ronnie b. wilbur and dietmar roehm: visual boundaries in sign motion. 175 https://doi.org/10.3765/elm https://www.elm-conference.net/ the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism nattanun chanchaochai* abstract. using the negated universal quantifier not every, the study investigates the interpretations of scalar implicatures, lexical presuppositions, and implicated presuppositions by thai children with autism spectrum disorders (asds; n = 32), compared to their typically-developing (td) peers (n = 70) and adults (n = 40). the results provide further empirical evidence to the literature (chevallier et al. 2010, hochstein et al. 2017, pijnacker et al. 2009) that not only do adolescents with asd perform on par with td adolescents, children with asd are also age-appropriate in their performance on deriving scalar implicatures. despite the children with asd’s ability to compute scalar implicatures, they still tend to give more logical, literal responses, compared to their peers. compared to adults, both children with asd and td children still have a higher tendency to rely on the logical meaning rather than pragmatically inferred meaning. no additive effects of implicated presuppositions are found in any group of the participants. keywords. scalar implicatures; presuppositions; implicated presupposition; acquistion; typically-developing children; children with autism spectrum dirorder 1. introduction. autism spectrum disorder (asd) is a developmental disorder defined by a dyad of impairments, including social communication impairments and restricted, repetitive patterns of behaviors and interests (american psychiatric association 2013). pragmatic and discourse deficits have long been accepted to be central to the characteristics of asd. owing to the prevalence of pragmatic deficits across the spectrum, this domain has been the focal point of research for the past several decades (kanner 1943, tager-flusberg 1999; a.o.). while the vast majority of literature on pragmatics and autism based its conclusions – that children with asd have pragmatic deficits – solely on the socially or contextually-dependent, less linguistically-informed side of pragmatics, first attempts on studying the linguistically-associated side found no difficulties for adolescents with asd and adults on tasks involving scalar implicature (chevallier et al. 2010, hochstein et al. 2017, pijnacker et al. 2009). these results suggest that certain less-explored parts of pragmatics in autism may still be intact. the interplay between semantics and pragmatics plays a crucial role in the understanding of a language. meaning in language is not always lexically encoded or grammatically derived. in addition to the literal, truth-conditional meaning, an utterance may also have other contextually influenced meanings. while the meaning inferred through particularized conversational implicatures (pcis; grice 1975, levinson 1983, 2000) arises only by virtue of a particular conversation, not from any lexical or grammatical components of the utterance, the meaning implicated through *this article is based on a part of my publicly defended phd dissertation at the university of pennsylvania. i would like to express my deepest gratitude to florian schwarz for his guidance and valuable advice. i would like to also thank david embick, kathryn schuler, jeremy zehr, muffy siegel, other members of the experimental study of meaning lab at upenn, and the audience at the first experiments in linguistic meaning (elm1) conference. author: nattanun chanchaochai, chulalongkorn university (nattanun.c@chula.ac.th). proceedings of elm 1: 101-112, 2021 c©2021 nattanun chanchaochai published by the lsa with permission of the author(s) under a cc by license. 101 https://doi.org/10.3765/elm https://www.elm-conference.net/ generalized conversational implicatures (gcis) is still linguistically tied. the latter type of pragmatic meaning, however, is still cancelable and not a part of the inherent, semantic meaning. presuppositions are another important type of inference in language, allowing more than one proposition to be communicated in one single sentence. they also serve as an indication of which proposition is the main assertion and which is merely background information (sauerland 2008a). while the extent to which presuppositional inferences are semantically or pragmatically driven is still the subject of considerable debate, presuppositions are typically assumed to convey the information that is already known but taken for granted by the speakers and still be projected under embedding operators, unlike the literal content that can be canceled by these operators (chierchia & mcconnell-ginet 1990, karttunen 1973; a.o.). additionally, heim (1991) observed that the infelicities of certain expressions cannot be accounted for by either lexical presuppositions or conversational implicatures. she proposed the maximize presupposition maxim, raising the possibility that implicated meanings could also be at play at the level of presuppositional information. her maxim suggests that the form with the strongest lexical presupposition must be chosen whenever its presupposition is felicitous. in other words, an utterance should lexically presuppose as much as possible. the idea of implicated presupposition (sauerland 2003, 2008a,b) is that in the case where the lexical entry with the strongest lexical presuppositions is not chosen, it can be implicated that the left-out presuppositions are not assumed. while it is claimed that implicated presuppositions are pragmatically derived in the same fashion as implicature, they are still backgrounded in the presuppositional domain. this hybrid type of inference between implicature and presupposition is thus interesting in its theoretical stance and in an acquisition point of view. by exploring both the linguistically-informed and contextually-informed meaning of an utterance, this paper provides insights on what aspects of meaning in language are particularly difficult for children with asd. the present paper is structured as follows. section 2 provides the background on previous experimental studies on quantifiers. section 3 then proceeds to present the methods of our experiment. the results are provided in section 4 and discussed in section 5. 2. previous experimental studies. this section reviews the relevant literature on semantic and pragmatic inferences of the quantifier ‘every’ (section 2.1), previous experimental studies on quantifiers in child language (section 2.2), and on quantifiers and autism (section 2.3). 2.1. semantic and pragmatic inferences of the quantifier ‘every’. yatsushiro (2008) viewed that the oddness of the following utterances in (1) are due to the fact that they violate the three presuppositions of the universal quantifier ‘every’ in (2), related to the set of its first argument. (1) a. #every tail of mine is long and curly. b. #every tongue of mine is pink. c. #every leg of mine is muscular. (2) a. existential presupposition: there exists at least one member. b. anti-uniqueness presupposition: there exists more than one member. c. anti-duality presupposition: there exists more than two members. while existential presupposition is a part of the lexical meaning of the quantifier ‘every’, the other two are implicated presuppositions, whereby there is an expression with a stronger presupproceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 102 https://doi.org/10.3765/elm https://www.elm-conference.net/ position, e.g., ‘the’ for (2-b) and ‘both’ for (2-c), for a speaker to use. speakers are required to choose the term that has the strongest presupposition, abiding by the pragmatic maxim of maximize presupposition (heim 1991). hence, it is more felicitous to say the following sentences in (3) than the previous (1). (3) a. my tongue/the tongue of mine is pink. b. both of my legs are muscular. quantifiers, such as ‘every’, may also involve scalar implicatures. orderings on lexical scales express scalar relations between scalar alternatives within the same set, based on their semantic strength defined by entailment. a given lexical item outranks, i.e., is stronger than, its alternate on the same scale if and only if a statement with its presence unidirectionally entails the corresponding statement containing its alternate (horn 1972). scalar implicatures may hold in the case of numbers on numeric scales, where the situations that are compatible with an utterance described with a larger number are a superset of the situations that are compatible with an utterance with a smaller number. similarly, assuming that some and all are scalar alternatives, the use of the expression some, which is ordered lower on the scale of quantity than the expression all, implicates that the use of all is not applicable, as seen that the utterance (4-a) typically implicates (4-b). (4) a. some graduate students have finished writing their term paper. b. not all graduate students have finished writing their term paper. 2.2. quantifiers in child language. many studies have investigated children’s comprehension of concepts related to quantification, including approximation, numbers, sets, and quantifiers (see lidz (2016) and smits (2010) for a comprehensive review). noveck (2001) recruited 8-yearolds, 10-year-olds, and adult native speakers of french to participate in the study. compared to the adults, the children in both groups were significantly more accepting of sentences such as ‘some giraffes have long necks’, suggesting that children prefer more logical responses than adults. some subsequent studies yielded similar results as noveck (2001) (see gualmini et al. 2001, foppolo et al. 2012; a.o.), with one major note that the children’s dispreference for calculating scalar implicatures may not arise from their genuine inability to do so but may be due to experimental settings. papafragou & musolino (2003) found that 5-year-old greek-speaking children were highly significantly more likely than adults to judge pragmatically infelicitous descriptions as being true. however, when they adjusted the experimental procedures and provided some training to another group of 5-year-olds, they observed a significantly higher rejection rates than those in the first version of experiment, although the children still did not reach the adult-like levels. as laid out earlier, the universal quantifier ‘every’ may also involve inferences from lexical and implicated presuppositions. yatsushiro (2008) reported that 6-year-old german-speaking children were more likely than adults to accept an equivalent sentence of ‘every girl here is playing soccer’ even when the picture they were shown depicted only one girl playing soccer. this suggests that children may base their felicity judgement merely on the lexical existential presupposition. the results support that lexical presuppositions are acquired earlier than implicated presuppositions of anti-uniqueness, which asks for more than one member in the set. moreover, yatsushiro (2008) found further, although less concrete, evidence for her hypothesis that the acquisition of implicated presuppositions and scalar implicatures may pattern together in their path. proceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 103 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.3. quantifiers and autism. first attempts at studying the linguistically-associated side found no difficulties for adolescents with asd and adults on tasks involving scalar implicature. pijnacker et al. (2009) found no statistical difference between high-functioning asd and td controls in their responses in judging underinformative some sentences such as ‘some sparrows are birds’ and underinformative disjunction sentences such as ‘zebras have black or white stripes’ to be false. chevallier et al. (2010) followed up on pijnacker et al. (2009), using spoken language stimuli with added stress on the disjunction or. contrary to their prediction, they replicated the unexpected results in pijnacker et al. (2009) that adolescents with asd equally accept pragmatically inferred disjunction as being correct, compared to the td group. hochstein et al. (2017) also reported similar results in adolescents with asd (n = 18; age m = 14.9, range = 12–18), compared to neurotypical adults (n = 17; age m = 22.6; range = 18–41). however, they found that adolescents with asd over-computed scalar implicatures in a different task where the participants need to base their answer on another person’s epistemic state. 2.3.1. the present study. this study employs the negated universal quantifier not every to more directly compare between scalar implicature and the two types of presupposition within one paradigm. similar to the universal quantifier every, the negated quantifier not every also yields an existential presupposition and an anti-uniqueness implicated presupposition. additionally, the literal none meaning is also present. given that if there is no intersection between restrictor and nuclear scope, cf. subject and predicate, other quantifiers, such as no or none, that are scalar alternatives to not every, would have been used to obey with the maxim of quantity, the meaning that there has to be a restrictor-nuclear scope intersection is then derived from the use of not every through scalar implicature. the four types of meanings for ‘not every’, derived from different mechanisms, are provided with example (5) below. (5) ‘not every boy is holding an ice cream.’ a. ∃ps (dom): existential presupposition for the domain there is a boy. b. ∃imp (restr ∩ scope): restrictor-nuclear scope intersection implicature there is a boy holding an ice cream. c. >1 impps (dom): anti-uniqueness implicated presupposition there is more than one boy. d. ¬∀: literal none meaning it is not true that every boy is holding an ice cream. 3. methods and design. this study adapted the covered box paradigm (huang et al. 2013). in each trial, a context picture was first shown on a screen to the participants (see the top picture in figure 1, with an auditory description in (6). the context picture depicts a group of animals doing the same thing, corresponding with the auditory context sentence. after the sentence ended, the screen was shifted to presenting two pictures, one visible and one covered with a black box, hidden from their view (see the bottom picture in figure 1 for illustration). note that in the visible pictures, the number of animals by type matched between the context screen and the test screen. another description was then auditorily presented with the scheme provided in (7). the participants were instructed to choose either the visible picture or the covered box as matching with the auditory proceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 104 https://doi.org/10.3765/elm https://www.elm-conference.net/ description. prior to the critical trials, four practice trials were presented. in the practice trials, the black box was removed to reveal the picture behind the box. in the critical trials, on the other hand, no picture behind the covered box was revealed. the participants made their choice by selecting the picture that they thought was a good match with the sentence in (7). (6) naj in klùm group ńı: this sàt animal thúk every tu:a cls y y jù: cont ... ... ‘in this group, every animal is doing y...’ (7) ... ... tè:-wâ: but naj in klùm group ńı: this x x mâj not thúk every tu:a cls y y jù: cont ‘but in this group, not every x is doing y’ [in this group, every animal is holding an ice cream, ...] (audio) [...but in this group, not every zebra is holding an ice cream.] (audio) figure 1: example context screen (top) and test screen (bottom). the conditions were manipulated with regards to the compatibility between the visible picture and the readings under investigation. to test how each of the four meanings, presented earlier in (5), plays a role in the participants’ interpretation of the quantifier not every, four experimental conditions were created. all of the critical experimental conditions are consistent with the literal none (¬∀) reading. they differ, however, in their consistency with the other three readings. table 1 presents example visible pictures in the test screens for each experimental condition and summarizes predictions of compatibility with the interpretations under investigation for each condition. the study consisted of 64 critical trials (16 trials per condition). in addition, 48 filler trials with the quantifier some (3 conditions; 16 trials per condition; see figure 2) were included to control for participants’ understanding of the task. those trials were counterbalanced and pseudo-randomized across 4 experimental lists. each list, therefore, contained 16 critical trials and 12 filler trials, presented with an even distribution of trial types in each of the four blocks. experimental lists were counterbalanced between participants. 3.0.1. procedure and subgroups of participants. a total of 92 children in the elementary grades were recruited from kasetsart university laboratory school, center for educational research and development. there were 32 children with asd (3 female; m age = 9;8; m nviq = 95.5) and 60 td children (11 female; m age = 7;11; m nviq = 116.6). all the participants with asd were classified in their medical records as having autistic disorder (ad). children with proceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 105 https://doi.org/10.3765/elm https://www.elm-conference.net/ ∃ps(dom) ∃imp(restr∩ scope) >1 impps (dom) ¬∀ allviolated * * * x impviolated x * x x impimppsviolated x * * x allmet x x x x note: xindicates a visible picture selection; * indicated a covered box selection. table 1: predictions for compatibility of readings in different conditions figure 2: filler conditions, shown with the auditory description “...only some animals...”. asd were matched to their td peers of the same gender in the same class. all of the participants had normal hearing and normal or corrected-to-normal vision. the studies were approved by the institutional review board of the university of pennsylvania. the parents of all the children provided written consent for them to participate in the study. the children and their parents were informed of their rights to withdraw from the study at any time. the children received toys and school supplies as compensation. the child data were collected offline. an additional collection of adult data was done online using penncontroller (zehr & schwarz 2018). the adult participants are native speakers of thai, demographically mixed, recruited through personal contacts and word-of-mouth. they accessed the experiment online using their personal computer. the consent form was shown at the beginning of the experiment. participants provided consent by their completion of the experiment. in the online version, participants were instructed to press f on their keyboard to select the visible picture or j to select the covered box. in the offline version, the participants indicated their choices by pointing. the only difference between the online and offline version of the task is that in the online version, before the experimental trials, there was a line of text on the screen emphasizing that their task is to match the sound with the second set of pictures on the screen in each trial. to ensure objective measures on whether a participant was performing the task, only the participants whose accuracy score was over 50% in the allmet condition, the filler 3 condition, or on proceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 106 https://doi.org/10.3765/elm https://www.elm-conference.net/ average between allmet and filler 3 were included in the statistical analyses. the two conditions that the criteria are based on are not controversial in their interpretation with respect to their corresponding auditory sentences. these criteria were also used to subgroup children into group 1 and group 2 based on their performance. three adult participants did not meet the criteria and, therefore, were excluded from the statistical analysis. a higher proportion of children in each participant group was included in the performance group 2 for this task. table 3.0.1 provides demographic details for each of the participant groups. adults asd td included excluded group 1 group 2 group 1 group 2 n (female n) 37 (15) 3 (1) 19 (2) 13 (1) 32 (7) 28 (4) m age 32.43 37 10;0 9;1 8;5 7;5 m nviq na na 99.79 89.24 120.71 111.97 table 2: participant information 3.0.2. predictions. this experiment aimed at exploring the acquisition of presuppositions, implicated presuppositions, and scalar implicature. the readings associated with ‘not every’ are laid out in (5) to gauge which reading is available for the participants’ interpretation of the negated quantifier. these readings are associated with a predicted response pattern across the experimental conditions, summarized in table 1. the compatibility with each reading was assessed to create unique contrasts across conditions. the ∃ps (dom) presupposition reading (5-a) predicts a contrast between the impimppsviolated (choice of visible picture) and the allviolated (covered) conditions. the ∃imp (restr ∩ scope) scalar implicature reading (5-b) predicts a contrast between the allmet (visible) and the impviolated (covered) conditions. the >1 impps (dom) implicated presupposition reading (5-c) predicts a contrast between the impviolated (visible) and the impimppsviolated (covered) conditions. it is worth noting, however, that the difference between the impimppsviolated and the impviolated conditions would be due to an additive and not pure effects from implicated presuppositions alone, as scalar implicatures are also violated in the condition. additionally, the pure literal ¬∀ reading (5-d) predicts a contrast between different groups of participants if they accept the allviolated visible picture to a different extent. 3.0.3. modeling. covered box rates were modelled separately for adults and children in the performance group 1. the child model contained 5 fixed effects factors, including condition (allviolated, impviolated, impimppsviolated, and allmet), participant group, z-scored ravens nonverbal iq, z-scored age, and gender. the model additionally include interactions between condition and participant group. the adult model contained 3 fixed factors, including condition, zscored age, and gender and no interaction. both models had a random effects factor for individual participants. dummy coding was employed for baseline re-levelling. additionally, to compare the results between children and adults, one additional model was fitted to the covered box rates, using 3 fixed effects factors of condition, participant group, and gender. interactions between condition and participant group were included, with a random effects factor for individual subjects. this model has to drop 2 fixed effects factors of age and nviq because no nviq data were collected for adults and age correlates with participant groups. proceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 107 https://doi.org/10.3765/elm https://www.elm-conference.net/ mixed effects logistic regression models were run using the lme4 package (version 1.1.12; bates et al. 2015) with the extension lmertest package (version 2.0.32; for obtaining p-values; kuznetsova et al. 2016) in the r software (version 3.3.1; r core team 2016) and mumin package (version 1.43.6; for obtaining r-squared; barton 2019). 4. results. a mixed effects logistic regression model was fitted to the covered box responses of the adult participants. covered responses were chosen significantly more in the allviolated condition, compared to the other three conditions, including the impviolated condition (b=-1.238, p=0.02), the impimppsviolated condition (b=-1.866, p<0.001), and the allmet condition (b=9.048, p<0.001). at the same time, the participants chose the visible picture to a significantly higher extent in the allmet condition than the other conditions: the impviolated condition (b=7.810, p<0.001) and the impimppsviolated condition (b=-7.182, p<0.001). the impviolated and the impimppsviolated conditions turned out to not differ from each other in their covered box rates (b=-0.628, p=0.119). additionally, on average of all conditions, female participants were found to accept the visible picture to a significantly higher rate than men (b=1.304, p=0.05). according to the mixed effects logistic regression model on the group 1 child data, children with asd and the td children in group 1 displayed the same pattern of significance contrasts between conditions as adults. in particular, the children with asd selected the covered box significantly more in the allviolated condition, compared to the other three conditions, including the impviolated condition (b=-1.731, p<0.001), the impimppsviolated condition (b=-2.411, p<0.001), and the allmet condition (b=-3.960, p<0.001). the visible picture was also chosen significantly more in the allmet condition than the impviolated condition (b=-2.228, p<0.001) and the impimppsviolated condition (b=-1.549, p=0.001). the impviolated and the impimppsviolated conditions appeared to be similar in their covered box rates (b=-0.679, p=0.082). similarly, the td children exhibited significantly stronger preference for the covered box in the allviolated condition than the impviolated condition (b=-2.614, p<0.001), the impimppsviolated condition (b=-2.932, p<0.001), and the allmet condition (b=-5.041, p<0.001). the allmet condition, on the other hand, significantly differed in its responses from the impviolated condition (b=-2.428, p<0.001) and the impimppsviolated condition (b=-2.109, p=0.001). the impviolated and the impimppsviolated conditions also did not differ in the td group (b=-0.319, p=0.290). additionally, group difference between the children with asd and the td children lies in their covered box rates in the allviolated condition, with the children with asd choosing significantly fewer covered responses (b=1.561, p=0.02). moreover, on average for all the children in the performance group 1, age (b=0.682, p<0.01), but not nviq (b=0.076, p=0.59), significantly increases their covered responses. figure 3 and 4 plotted mean covered response rates in adults and children, respectively. significance levels from the two models presented above are also summarized in the two figures. to further compare between the adult and child data, a logistic mixed effects model was fitted to their data, dropping the fixed effects of nviq and age. in general, adults were significantly more likely than both groups of children to choose a covered box in all of the conditions, except in the allmet condition, where children chose a covered box significantly more than adults. comparing the differences between (1) the allviolated and the impimppsviolated conditions (compared to asd: b=-0.59, p=0.374; to td: b=-1.123, p=0.08) and (2) the impviolated and the impimppsviproceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 108 https://doi.org/10.3765/elm https://www.elm-conference.net/ 95.3% 88.5% 83.1% 2% * *** *** ns *** *** allviolated impviolated impimppsviolated allmet a du lts % o f c ho os in g co ve re d bo x figure 3: adults’ accuracy by condition 73.7% 43.4% 31.6% 11.8% *** *** *** ns *** ** 85.9% 50% 44.5% 14.1% *** *** *** ns *** *** asd td allviolated impviolatedimpimppsviolated allmet allviolated impviolatedimpimppsviolated allmet g ro up 1 c hi ld re n % o f c ho os in g co ve re d bo x figure 4: group 1 children’s accuracy by condition olated conditions (compared to asd: b=-0.073, p=0.90; to td: b=0.298, p=0.55), adults did not differ from either group of children. however, adults differ from both groups of children in their differences between the covered box rates in the impviolated and the allmet conditions (compared to asd: b=5.428, p<0.001; to td: b=5.269, p<0.001). no logistic regression model was run on the data of the children in the performance group 2. in general, the children in this group seemed to misunderstand or not fully comprehend the task, resulting in similar pattern of covered box rates across conditions (asd: 65.4% for allviolated, 75% for impviolated, 71.2% for impimppsviolated, and 63.5% for allmet; td: 89.3% for allviolated, 81.2% for impviolated, 82.1% for impimppsviolated, and 90.2% for allmet). proceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 109 https://doi.org/10.3765/elm https://www.elm-conference.net/ 5. discussion. the adult data provide empirical evidence that they were accessing the meanings from the existential lexical presupposition and the restrictor-nuclear scope intersection implicature, as suggested from their significant differences in covered box rates between (1) the allviolated and the impimppsviolated conditions and (2) the impviolated and the allmet conditions, respectively. the adults, however, did not display their sensitivity to the anti-uniqueness implicated presuppositions, showing no differences between the impimppsviolated and the impviolated conditions. the pattern in the child data was the same as adults, with signs of them computing lexical presuppositions and scalar implicatures, but not implicated presuppositions. overall, age was a significant factor in predicting their responses, with older children choosing more covered boxes. while the overall patterns are similar, the children with asd chose covered boxes significantly less in the allviolated condition, compared to the td children. other pairs of conditions did not yield such differences. this suggests that access of the literal, logical meaning is the most indicative of group differences in children, with children with asd basing their interpretation on literal meanings to a higher extent than td children. in contrast, children with asd are on par with td children in accessing the meaning derived from lexical presuppositions, implicated presuppositions, and scalar implicatures. a comparison between adults’ and children’s behavior reveals several significant differences. for one thing, adults had significantly higher covered box rates in the allviolated, the impviolated, and the impimppsviolated conditions, while having significantly lower covered box rates in the allmet condition than both groups of children, suggesting that they were, in general, more likely to produce expected results. secondly, adults were significantly less likely to accept the visible picture in the allviolated condition than both children with asd and td children, indicating the children’s higher tendency to rely on the logical meaning rather than pragmatically inferred meaning, compared to adults. additionally, adults significantly chose more covered boxes in the impviolated condition than in the allmet condition to a higher extent than children in both groups. this suggests that children are less likely to derive scalar implicatures, compared to adults. the adults and both groups of the children, however, are not different in their likelihood to derive lexical presuppositions. the additive effect of implicated presupposition violation is also not greater in the adults, compared to the groups of children. overall, the results are very consistent with previous literature in many aspects. firstly, previous experimental studies on scalar implicature in child language found significantly higher preference for logical responses to scalar implicature in children than adults (gualmini et al. 2001, foppolo et al. 2012, noveck 2001, papafragou & musolino 2003; a.o.). the current study also observed that children rely more on logical, literal meaning to a higher extent compared to adults. secondly, studies on scalar implicatures in adolescents with asd suggest that their scalar implicatures are intact (chevallier et al. 2010, hochstein et al. 2017, pijnacker et al. 2009). the results of this study provide further empirical evidence to the literature that not only do adolescents with asd perform on par with td adolescents, children with asd are also age-appropriate in their performance on deriving scalar implicatures. additionally, this study adds that even though the children with asd’s ability to compute scalar implicature is on par with td children, they still tend to give more logical, literal responses, compared to their peers, as seen in their significantly lower covered box rates in the allviolated condition, and not in other conditions. proceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 110 https://doi.org/10.3765/elm https://www.elm-conference.net/ thirdly, the current results are in accordance with yatsushiro (2008)’s claim that lexical presuppositions are acquired earlier than scalar implicatures and perhaps than certain types of implicated presuppositions, similar to legendre et al. (2011). the current results also adds to the literature that types of implicated presupposition matter in the acquisition pattern. this is evident from the fact that even though implicated presuppositions seem to affect the accuracy rates in comprehending certain personal reference terms (chanchaochai 2017), they do not seem to have an additive effect in this study of the negated quantifier. the proposal that different types of implicated presuppositions may affect participants’ performance differently is a plausible proposal, considering that rates of deriving scalar implicatures were also previously observed to differ by scalar terms. papafragou & musolino (2003) improved success rates in children deriving scalar implicatures when the task involved number terms, such as , rather than scalar terms, such as . these observations raise interesting theoretical issues on types of implicated presupposition and their pattern of acquisition. further investigation into the acquisition of different types implicated presuppositions, compared to other types of pragmatic inferences, should be made, probing the pure, not additive, effects from implicated presuppositions. references american psychiatric association. 2013. diagnostic and statistical manual of mental disorders (5th ed.). arlington, va: american psychiatric association. barton, kamil. 2019. mumin: multi-model inference. https://cran.r-project.org/ package=mumin. r package version 1.43.6. bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67(1). 1–48. 10.18637/jss.v067.i01. chanchaochai, nattanun. 2017. on acquiring a complex personal reference system: experimental results from thai children with autism. poster presented at sinn und bedeutung 22, september 7-10. chevallier, coralie, deirdre wilson, francesca happé & ira noveck. 2010. scalar inferences in autism spectrum disorder. journal of autism and developmental disorders 40. 1104–1117. chierchia, gennaro & sally mcconnell-ginet (eds.). 1990. meaning and grammar. an introduction to semantics. cambridge: mit press. foppolo, francesca, maria teresa guasti & gennaro chierchia. 2012. scalar implicatures in child language: give children a chance. language learning and development 8. 365–394. grice, herbert paul. 1975. logic and conversation. in peter cole & jerry l. morgan (eds.), syntax and semantics: speech acts, vol. 3, 41–58. new york: academic press. gualmini, andrea, stephen crain, luisa meroni, gennaro chierchia & maria teresa guasti. 2001. at the semantics/pragmatics interface in child language. in rachel hastings, brendan jacksonand & zsofia zvolenszkyn (eds.), proceedings of the 11th semantics and linguistic theory (salt11) conference, 231–247. ithaca, ny: clc-publications, cornell university. heim, irene. 1991. artikel und definitheit. in von a. stechow & d. wunderlich (eds.), semantik: ein internationales handbuch der zeitgenoessischen forschung, 487–535. berlin: de gruyter. hochstein, lara, alan bale & david barner. 2017. scalar implicature in absence of epistemic reasoning? the case of autism spectrum disorder. language learning and development . proceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 111 https://doi.org/10.3765/elm https://www.elm-conference.net/ horn, laurence r. 1972. on the semantic properties of logical operators in english: department of linguistics, university of california, los angeles doctoral dissertation. huang, yi ting, elizabeth spelke & jesse snedeker. 2013. what exactly do numbers mean? language learning and development 9. 105–129. kanner, leo. 1943. autistic disturbances of affective contact. nervous child 2. 217–250. karttunen, lauri. 1973. presuppositions of compound sentences. linguistic inquiry 4. 169–193. kuznetsova, alexandra, per bruun brockhoff & rune haubo bojesen christensen. 2016. lmertest: tests in linear mixed effects models. https://cran.r-project.org/package= lmertest. r package version 2.0-32. legendre, géraldine, isabelle barrière, louise goyet & thierry nazzi. 2011. quantifier acquisition: presuppositions of ‘every’. in mihaela pirvulescu et al. (eds.), selected proceedings of the 4th conference on generative approaches to language acquisition north america (galana 2010), 150–162. somerville, ma: cascadilla proceedings project. levinson, stephen c. 1983. pragmatics. cambridge: cambridge university press. levinson, stephen c. (ed.). 2000. presumptive meanings: the theory of generalized conversational implicature. cambridge: mit press. lidz, jeffrey l. 2016. quantification in child language. in jeffrey l. lidz, william snyder & joe pater (eds.), the oxford handbook of developmental linguistics, 498–519. oxford: oxford university press. noveck, ira a. 2001. when children are more logical than adults: experimental investigations of scalar implicature. cognition 78. 165–188. papafragou, anna & julien musolino. 2003. scalar implicatures: experiments at the semanticspragmatics interface. cognition 86. 253–282. pijnacker, judith, peter hagoort, jan buitelaar, jan-pieter teunisse & bart geurts. 2009. pragmatic inferences in high-functioning adults with autism and asperger syndrome. journal of autism and developmental disorders 39. 607–618. r core team. 2016. r: a language and environment for statistical computing. r foundation for statistical computing vienna, austria. https://www.r-project.org/. sauerland, uli. 2003. a new semantics for number. in proceedings of salt13, 258–275. sauerland, uli. 2008a. implicated presuppositions. in anita steube (ed.), sentence and context, mouton de gruyter. sauerland, uli. 2008b. on the semantic markedness of phi-features. in daniel harbour, david adger & susana bejar (eds.), phi theory: phi features across interfaces and modules, 57– 82. oxford university press. smits, erik-jan. 2010. acquiring quantification: how children use semantics and pragmatics to constrain meaning: department of linguistics, university of groningen doctoral dissertation. tager-flusberg, helen. 1999. a psychological approach to understanding the social and language impairments in autism. international review of psychiatry 11. 325–334. yatsushiro, kazuko. 2008. quantifier acquisition: presuppositions of ‘every’. in proceedings of sinn und bedeutung 12, 663–677. oslo: university of oslo. zehr, jérémy & florian schwarz. 2018. penncontroller for internet based experiments (ibex) 10.17605/osf.io. https://doi.org/10.17605/osf.io/md832. proceedings of elm 1: 101-112, 2021 nattanun chanchaochai: the interpretations of scalar implicatures, presuppositions, and implicated presupposition by thai children with autism. 112 https://doi.org/10.3765/elm https://www.elm-conference.net/ the enduring effects of default focus in let alone ellipsis: evidence from pupillometry. jesse a. harris∗ abstract. the study of clausal ellipsis in sentence processing has revealed that comprehenders are sensitive to multiple, sometimes conflicting, pressures when recovering elided content. this paper presents a pupillometry experiment investigating how the human language processing system responds to sentences in which the location of a pitch accent clashes with global preferences for local correlates. the results are discussed in light of existing literature, including the enduring focus principle, in which locations for default pitch accent continue to influence focus-sensitive processes regardless of overt markers of focus. keywords. ellipsis; sentence processing; prosody; pupillometry 1. introduction. ellipsis poses a set of analytical and empirical problems for the study of the human language processing system. the processor must impute the elided content by establishing an interpretive link with an antecedent clause. as the interpretation of ellipsis depends on many factors, placing different kinds of information in conflict affords psycholinguists the unique opportunity to investigate aspects of the underlying architecture of the language processing system. in contrastive clausal ellipsis, the remnant is placed in focal contrast with its correlate. a particularly intriguing case is let alone ellipsis, as in john can’t run a mile, let alone a marathon (hulsey 2008, toosarvandani 2010, harris 2016, to appear, for ellipsis analyses). to interpret the remnant (a marathon), the processor locates the contrasting correlate phrase (a mile) in the antecedent clause from among other same-category competitors using multiple, possibly competing, preferences. experimental and corpus research finds that the nearest / most local possible correlate is vastly preferred (harris & carlson 2016, 2018). similar biases have been observed for other clausal ellipsis structures, like sluicing (frazier & clifton 1998) and replacives (carlson 2002). however, semantic and prosodic parallelism have also been shown to interact with locality (harris & carlson 2016, 2018), suggesting a general, but violable, preference for pairing a remnant with a correlate that is maximally similar along multiple dimensions. this paper concentrates on how contrastive pitch accent location interacts with global preferences for local correlates in the let alone construction. the pupillometry method is employed as a means to identify processing costs during online auditory comprehension. an introduction to let alone ellipsis is provided next, followed by a brief discussion of pupillometry. it is argued that misleading pitch accent asymmetrically interrupts the processing of let alone ellipsis, providing partial support for the enduring focus principle (harris & carlson 2018). 2. the let alone structure. sentences with let alone exhibit a complex interaction of syntactic, semantic, pragmatic, and prosodic properties (fillmore, kay & o’connor 1988, hulsey 2008, ∗many thanks to the amazing undergraduate research assistants in the ucla processing lab for adminstering the study, including ani babekhanean, lily kawaoto, joonhwa kim, alison suh, marina suh, chenchen wang, richard wang, and rebecca wu. thanks also to the audiences at elm2 and wccfl 40 for helpful feedback. author: jesse a. harris, university of california los angeles (jharris@humnet.ucla.edu). proceedings of elm 2: 117-128, 2023 c©2023 jesse a. harris published by the lsa with permission of the author(s) under a cc by license. 117 https://doi.org/10.3765/elm https://www.elm-conference.net/ toosarvandani 2010). central properties are provided in (1).1 (1) properties of let alone a. coordinates or compares elements in contrastive focus; b. presupposes a scalar relationship between items in contrastive focus along some contextually salient scale; c. typically licensed by negative element; d. hosts stripping ellipsis. for example, speaker b might answer a’s question using let alone in (2). the elements under comparison are marked with contrastive pitch accent (denoted by small caps). these items are placed on a contextually salient scale that might be reconstructed as something like eating escargot is less likely (or less desirable, etc.) than eating caviar. negating a lower element on the scale contextually entails the negation of any higher element (toosarvandani 2010). (2) a. did john try the any of the special hors d’oeuvres tonight? b. john didn’t eat the caviar correlate , let alone escargot remnant . a stripping ellipsis approach will be adopted for such cases (harris 2016, to appear). the fragment (escargot) following let alone corresponds to the remnant of ellipsis, and is paired with a correlate (caviar) in the antecedent clause. assuming a move-and-delete style analysis, the remnant fronts into a focus (3-a) or contrastive topic (3-b) position, while the given material in the rest of the clause is deleted (denoted by 〈·〉).2 (3) a. john didn’t eat caviar, let alone [focp escargot]1 〈john eat t1〉 b. john didn’t eat caviar, let alone [ctp sue]1 〈 t1 eat caviar〉 while object contrasts (3-a) seem fairly natural, subject contrasts (3-b) are deemed less acceptable and appear only rarely in corpora (harris & carlson 2016, 2018). intuitively, the implausible parse in which sue is an object is at least temporarily available, despite the explicit pitch accent on the subject, producing an effect reminiscent of the garden path produced by coordination ambiguities (frazier 1987, engelhardt & ferreira 2010). the rest of this paper is devoted to exploring a series of related issues: (i) why ellipsis structures with subject contrasts might be perceived as less acceptable, (ii) how factors such as pitch accent placement contribute to the effect, and (iii) how pitch accent placement affects the online processing of let alone ellipsis. 3. processing ellipsis. the processing of elliptical sentences is complex. it is likely that the sentence processing system draws on many different information channels to establish an interpretation of the ellipsis site. in the case of clausal ellipsis, the interpretation requires the recognition of the ellipsis site, some kind of representation of the elided material, and the location of a correlate 1note that some dialects permit a positive use of let alone (fillmore et al. 1988, toosarvandani 2009), which does not require negation and presents an afterthought, rather than a scalar comparison between items in contrastive focus (cappelle, dugas & tobin 2015). such cases are relatively rare in corpora (harris & carlson 2016). 2for purposes of exposition, i will adopt a syntactic approach to ellipsis, although few, if any, of the central points below depend on this conception. proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 118 https://doi.org/10.3765/elm https://www.elm-conference.net/ for the remnant (yoshida 2018). in particular, i follow harris & carlson (2018) and assume that processing clausal ellipsis requires, at a minimum, the processor to engage in the following tasks: (4) basic tasks of the processor in ellipsis processing: a. parse the remnant by constructing the appropriate phrase structure for the remnant given the input. b. locate the correlate, if any, from the antecedent clause. c. construct the elided phrase by regenerating or copying a structure at logical form. task (4-a) falls under the domain of standard parsing routines, in which phrase structure is assigned to a string. tasks (4-b) and (4-c) have been the focus of much previous research. in what follows, our attention will remain on the former, and so discussion of the latter will consequently be suppressed (however, see frazier 2018, yoshida 2018, for review). the second task will be referred to as the remnant-correlate pairing process. recent studies have concentrated on how the following two major factors guide remnantcorrelate pairing: (5) recency/locality: prefer the object / closest correlate (e.g., frazier & clifton 1998, harris & carlson 2016). (6) parallelism: match internal properties (e.g., semantics, pitch accent) of dps in similar structural positions (e.g., carlson, 2002). evidence for a preference for recent or local correlates has been observed in many different types of clausal ellipsis, including sluicing (frazier & clifton 1998, carlson et al. 2009, harris 2015, 2019), replacives (carlson 2002, 2013), and let alone ellipsis (harris & carlson 2016, 2018). to take the case of let alone ellipsis, harris & carlson (2016) found that the vast majority (approximately 84%) of remnants were paired with the nearest correlate in corpora. in two self-paced reading studies, they also found an online cost for violating locality. for example, in one study, the location of a contrastive adjective like nicest was placed in either object (7-a) or subject (7-b) position. the adjective forms a parallel contrast with the remnant the meanest one, which was held constant across the two conditions. (7) a. the nurse couldn’t stand the nicest patient, let alone the meanest one, and no one at the hospital was happy at all. b. the nicest nurse couldn’t stand the patient, let alone the meanest one, and no one at the hospital was happy at all. sentences with a subject position adjective were read slower on the remnant and the region immediately following (and no one). the results suggest that associating the remnant with a non-local correlate taxes the processing system, despite clear semantic parallelism between the correlate and the remnant. another way to establish parallelism between the correlate and the remnant is through prosodic marking of information structure, which is integral to ellipsis in at least two ways: first, only given material can be elided and, second, the remnant to ellipsis must contain new or contrastive information (winkler 2018, for overview). in a corpus study of radio interviews, harris & carlson (2018) proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 119 https://doi.org/10.3765/elm https://www.elm-conference.net/ found that every correlate and remnant in let alone ellipsis bore pitch accent, most of which were realized as a l+h* contrastive accent (79% on correlates and 73% on remnants). in acceptability judgments, however, the location of pitch accent only partially mitigated the cost for non-local correlate-remnant pairings. in that study, ratings decreased for both sentences with either a subject pitch accent (8-b) or a subject remnant (let alone kayla). the factors interacted, so that subject pitch accent reduced the ratings cost of a subject remnant. (8) a. object accent: danielle didn’t pass the quiz, let alone # kayla / the final b. subject accent: danielle didn’t pass the quiz, let alone kayla / # the final crucially, however, mismatching cases elicited qualitatively different kinds of penalties. a subject remnant paired with an object accent elicited a greater cost than an object remnant paired with a subject accent. following suggestions by büring (2012, 2016), harris & carlson (2018) proposed the enduring focus principle, in which “locations that typically bear default focus continue to provide potential locations for focus, regardless of overt markers of focus.” in the context of clausal ellipsis, the object remains a tempting correlate for let alone ellipsis because the object receives nuclear pitch accent by default in english svo sentences. the failure of overt pitch accent to overturn locality preferences was attributed not to a specific preference for correlate location in the antecedent clause, but to global informational structure expectations for focus placement. in addition, listeners might be more willing to tolerate some degree of prosody-focus misalignment in auditory processing. according to enduring focus, listeners would thus be more tolerant of mismatches that preserved global expectations of the prosody. although there is evidence that sentences with a local correlate are advantaged in online processing during silent reading, whether locality operates during online auditory sentence processing remains an open question. recent research indicates that language comprehenders are strongly guided by their expectations of the input, and that those expectations are likely to be influenced by multiple factors including grammatical constraints, experience, contextual information, and what is known about the speaker. one possibility is that prosodic expectations for pitch accent location will guide auditory sentence processing, allowing listeners to accommodate or correct prosodic mismatches relatively quickly. in this case, the mismatch asymmetry would be replicated in online processing. another possibility is that the accommodation of prosodic mismatch is delayed, perhaps as a post-interpretive repair process. under this scenario, any penalties generated by mismatching pitch accent would be of similar magnitude regardless of correlate location. the experiment below capitalizes on the unique advantages of pupillometry to investigate these questions. as this technique is not widely used, a brief introduction to pupil dilation and its sensitivity to linguistic variables follows. after discussing the experiment and its central findings, the paper concludes first by contextualizing the results within the current literature on ellipsis and prosodic processing and then by exploring alternative accounts of the findings. 4. pupillometry study. 4.1. method. pupil size fluctuates naturally in response to multiple factors. in addition to environmental factors, such as change in luminance, the pupil dilates in response to increases in cognitive load, mental effort, and emotional stimulation (laeng, sirois & gredebäck 2012, for proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 120 https://doi.org/10.3765/elm https://www.elm-conference.net/ review). the pupil begins to respond within 200ms of stimulus onset, yet its cumulative effects are relatively slow, peaking approximately 700-1200ms after stimulus onset. pupillometry is the quantitative study of pupil size change over time. early pupillometry studies found that tasks that involved greater memory and attention (hess & polt 1960, 1964, kahneman & beatty 1966), including sentences that were difficult to process (just & carpenter 1993), were associated with greater pupillary excursions. after a period of relative quiet, pupillometry is now enjoying a resurgence in the psycholinguistics literature (schmidtke 2018). recent studies have found an association between increased pupil size and lexical frequency or increased emotional valence (kuchinke et al. 2007), structurally complex sentences (demberg & sayeed 2016), prosodic disambiguation in garden path sentences (engelhardt, ferreira & patsenko 2010, harris & jun 2018), attachment ambiguities (harris, lawn & kaps 2019a, harris et al. 2019b), semantic anomalies (demberg & sayeed 2016), inadequate or misleading pitch accent (zellin et al. 2011), and violations of expected meter (scheepers et al. 2013, breiss, harris & rysling 2021). while there is no perfect measurement for studying cognition, pupillometry is appealing for many reasons. relative to neurophysiological measures, it is inexpensive and easy to administer. further, pupil size cannot be controlled consciously and is unlikely to reflect task-specific strategies employed during the experiment. in addition, pupillometry is particularly useful in auditory sentence processing as the listenter may be required to engage in no other task besides listening naturally and for comprehension. thus, pupillometry studies offer a highly promising avenue for exploring the role of prosodic information in online sentence processing. 4.2. design and materials. the 20 items from experiment 1 in harris & carlson (2018) formed the basis of the 2×2 design, crossing remnant type (objectrem vs. subjectrem) with contrastive pitch accent location on the correlate (objectpa vs. subjectpa). experimental materials were modified to include post-remnant material on which to record pupil size dilation. items were then re-recorded by a trained phonologist familiar with tobi, but who was not involved in the project. materials were inspected for pitch accent location and overall quality and were re-recorded as needed. audio files were truncated after the remnant, where 100ms of computer generated silence was inserted. one critical post-remnant region was selected from each quartet and was spliced into each condition, so that the pupillary response was recorded on the identical acoustic material (11). the procedure was repeated for each quartet. the resulting recording was then normalized and checked for naturalness. (9) object remnant a. object accent: danielle didn’t pass the quiz, let alone the final b. subject accent: danielle didn’t pass the quiz, let alone final . . . (10) subject remnant a. object accent: danielle didn’t pass the quiz, let alone kayla . . . b. subject accent: danielle didn’t pass the quiz, let alone kayla . . . (11) follow on clause for pupil recording: before the evening class on business law. proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 121 https://doi.org/10.3765/elm https://www.elm-conference.net/ half of the items were followed by two-alternative forced choice comprehension questions (12). (12) comprehension question: what was the class on? a. law b. ethics filler items consisted of 40 sentences from two unrelated experiments on attachment ambiguity and 20 non-experimental fillers. half of the fillers were followed by comprehension questions similar to (12). 4.3. participants and procedure. forty-eight self-reported native speakers of english were recruited from ucla to participate in the study, which lasted 30 minutes on average. all subjects reported having normal hearing. subjects received course credit for their participation. materials were presented with experiment builder (sr research) over high-quality seinhauser sound-isolating headphones. eye position and pupil area were recorded using an sr research eyelink 1000 plus eye tracker sampling at 250hz. the tracker was mounted to the table at 55cm from a 27 inch lcd monitor with a light gray background. the room was moderately lit at consistent levels to control the effect of light contamination on pupil size. a 5-point calibration procedure was used before recording and as necessary, and a drift correction was conducted before every trial. subjects were instructed to avoid blinking while the sentence played. after each pre-trial drift correct, subjects were allowed to rest their eyes and blink as needed before self-initiating the trial with a gamepad. after listening to the sentence, subjects answered any comprehension question associated with the sentence by selecting their answer on a gamepad. pupil size was recorded for the entire sentence with a high-speed sr research eye-tracker. only the 2 second period after the offset of the remnant is reported here. in this period, material was acoustically identical within a quartet. data cleaning followed the recommendations in the literature (mathôt et al. 2018, winn et al. 2018, van rij et al. 2019). blinks and other artefacts were removed automatically with pupilpre (kyröläinen et al. 2019) with 200ms of padding around the event. any remaining blinks or track losses were removed by hand through visual inspection. the 100ms period of silence after the remnant was taken as the baseline period to calculate change over time. trials with less than 50% of data in the baseline or less than 80% of data in the period of interest were removed. subjects with less than 4 trials per condition were excluded. missing data points on the remaining data were interpolated with spline-smoothing. the data was downsampled to 10hz to reduce autocorrelation. the data was then normalized by trial to reflect change in pupil size over time by subtracting the mean pupil size obtained from the 100ms baseline, rather than absolute pupil size. finally, a low-pass butterworth filter was applied to reduce fast measurement noise that is unlikely to originate from a physiological source 4.4. results. a generalized additive mixed effects model (gamm) was used to capture changes in pupillary excursion (van rij et al. 2019). gamms are useful for capturing non-linear relationships that develop over time. gamms use smoothing splines, polynomial functions divided into continuous segments, to find the best-fitting non-parametric curves fitting the data (see baayen & linke 2020, for a general introduction for linguists). predictor variables consisted of the factors underlying the experimental manipulation, i.e., remnant type and pitch accent location, along with proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 122 https://doi.org/10.3765/elm https://www.elm-conference.net/ their interaction. to account for variation among individual subjects and items, random effects included subjects and items over time. assuming that greater pupil response signals that more cognitive resources were consumed during processing, we make several key predictions. first, object remnants should be associated with less pupil dilation than subject remnants. second, pitch accent on an object should elicit less pupil dilation than pitch accent on a subject. finally, the two factors should interact, so that the subject pitch accent mediates the effect of a subject remnant. within the expected interaction, there are two possibilities, as the design generated two conditions where the pitch accent location matched the location of pitch accent on the correlate, (9-a) and (10-b), and two conditions where they did not, (9-b) and (10-a). pitch accent location could completely reverse any subject remnant penalty, so that the matching conditions pattern together: subjectrem-subjectpa (10-b) and objectrem-objectpa (9-a) would both be associated with lower pupil dilation. if, however, the location of pitch accent only partially mitigates the subject remnant penalty, subjectrem-subjectpa (10-b) might instead pattern with the subjectrem-objectpa (10-a) condition. results of the gamm are summarized in table 1. there was a penalty for subject remnants and a penalty for accented subject correlates, as predicted. there was also an interaction between the factors, in which the effect of subject remnants was reduced when the subject in the antecedent clause was pitch accented. estimate std. error t-value p-estimate (intercept) -1.47 4.81 -0.31 0.76 subjectpa 5.79 0.75 7.78 < .001 subjectrem 2.62 0.75 3.51 < .01 subjectpa × subjectrem -5.40 0.75 -7.24 < .001 table 1: summary of the generalized mixed effects model. to explore the interaction more closely, marginal means of the model were estimated with emmeans (lenth 2022). in the object remnant condition, accented object correlates significantly reduced the pupil size compared to accented subject correlates (β̂ = −22.39, se = 2.10, t = −10.64, p < .0001). in the subject remnant condition, however, pitch accent location did not induce significantly different pupillary responses (β̂ = −0.79, se = 2.11, t = −0.372, p = 0.71), suggesting that the location of pitch accent mitigated, but did not reverse, the subject remnant penalty. the interactive effect is clearly visualized in the plots in figure 1 below. as expected, the condition with an object remnant and an object pitch accent was associated with the least amount of pupil rise. when an object remnant was preceded by a pitch accented subject, the greatest effect on pupil change was observed. pitch accent location in the subject remnant conditions, in contrast, appeared to have no effect. in other words, prosodic parallelism did affect the pupillary response, but failed to completely reverse the effect of locality. 4.5. discussion. the findings of this pupillometry study partially corroborate harris & carlson’s (2018) offline ratings study in an online method. both methods found an advantage for object remnants over subject remnants, in support of an real-time preference for local correlate-remnant proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 123 https://doi.org/10.3765/elm https://www.elm-conference.net/ −40 −20 0 20 0 500 1000 1500 2000 time (ms) c ha ng e fr om b as el in e av er ag e object pa object rem subject pa object rem subject pa subject rem object pa subject rem normalized change in pupil sizea −20 −10 0 10 object pa object rem subject pa object rem subject pa subject rem object pa subject rem m ea n ch an ge fr om b as el in e av er ag e mean change in pupil sizeb figure 1: change in pupil size obtained by subtracting points from the 100ms baseline of silence at each 100ms bin (left panel). the average pupillary response obtained over the entire 2000ms post-remnant interval (right panel). pairings. further, subject accent elicited a general penalty compared to object accent, in keeping with a preference for accent in the default position. however, while there was evidence for a pitch accent mismatch asymmetry, the direction of the interaction differed from previous findings. in harris & carlson (2018), the effect of mismatch was greater for subject remnants than object remnants. in the present experiment, the mismatch effect was eliminated for subject remnants. there are several ways that the differences between studies might be reconciled, two of which are considered here. first, subject remnants require forming a contrast with the subject correlate. that process might require greater cognitive reserve in general, thereby delaying the prosodic mismatch effect with subject remnants. assuming that different measures capture distinct stages of processing, the effect of prosodic mismatch for subject remnants only becomes apparent in later stages of interpretation, perhaps after the object remnant parse has been completely eliminated. a second possibility is that the licensing conditions for subject remnants were not completely met in the materials. in harris (to appear), i argue that subject remnants are low contrastive topics (cts), syntactically located between tp and vp. contrastive topics are typically licensed in contexts with a pair list answer, a partial answer, or when the speaker initiates a topic shift to another individual büring (2003, 2014). without an appropriate question under discussion (qud), the listener must also accommodate some kind of salient relationship between the subject correlate and the remnant.3 in addition, cts tend to appear in multiple-focus constructions: a ct followed 3given that a contextual question placing contrast on the subject induces a reading time penalty (harris & carlson 2016), it is unlikely that providing a qud with an explicit subject contrast is alone sufficient to salvage a sentence with a non-local subject remnant. proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 124 https://doi.org/10.3765/elm https://www.elm-conference.net/ by another focused item that together constitute a question-answer pair (van hoof 2003, wagner 2012). it is possible that processing was delayed as listeners monitored the input for another focused element, e.g., the final, in the gapped variation in (13). (13) danielle didn’t pass the quiz, let kayla – the final if correct, let alone ellipsis with subject correlates may be more complex than let alone with object correlates, in terms of syntax or information structure, unlike what has been proposed for replacive ellipsis, alex hit bob, not carl (carlson 2002, stolterfoht 2005). a comparison between kinds of contrastive ellipses would be a highly informative direction for future research. 5. conclusion. the results of a pupillometry study investigating the interpretation of let alone ellipsis was reported. the experiment was designed to determine how the language processor responds to sentences in which prosodic parallelism and global locality preferences conflict. a penalty for non-local correlate-remnant pairings was observed. although pitch accent placement on a non-correlate elicited a greater pupillary response in object remnant cases, it had no effect on the subject remnant cases. consistent with previous studies on let alone ellipsis and the literature on ellipsis in english at large, the pattern was interpreted as reflecting the prioritization of syntactic over prosodic information in the interpretation of ellipsis (van der burght et al. 2021). while pitch accent type and location clearly guides processing expectations, it would appear that the syntactic information has a more robust effect when it comes to interpreting ellipsis. references baayen, r. harald & maja linke. 2020. an introduction to the generalized additive model. in magali paquot & stefan th gries (eds.), a practical handbook of corpus linguistics, 563– 591. new york: springer. breiss, canaan, jesse a. harris & amanda rysling. 2021. the online advantage of repairing metrical structure: stress shift in pupillometry. in tecumseh fitch, claus lamm, helmut leder & kristin teßmar-raible (eds.), the proceedings of the 43rd annual meeting of the cognitive science society, 2883–2890. vienna, austria. burght, constantijn l. van der, angela d. friederici, tomás goucha & gesa hartwigsen. 2021. pitch accents create dissociable syntactic and semantic expectations during sentence processing. cognition 212. 104702. büring, daniel. 2003. on d-trees, beans, and b-accents. linguistics and philosophy 26(5). 511– 545. büring, daniel. 2012. focus and intonation. in gillian russell & delia graff fara (eds.), routledge companion to the philosophy of language, 103–115. abingdon: routledge. büring, daniel. 2014. (contrastive) topic. in caroline féry & shin ishihara (eds.), handbook of information structure, 64–85. oxford, uk: oxford university press. büring, daniel. 2016. intonation and meaning. oxford, uk: oxford university press. cappelle, bert, edwige dugas & vera tobin. 2015. an afterthought on let alone. journal of pragmatics 80. 70 – 85. carlson, katy. 2002. parallelism and prosody in the processing of ellipsis sentences. new york: psychology press. proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 125 https://doi.org/10.3765/elm https://www.elm-conference.net/ carlson, katy. 2013. the role of only in contrasts in and out of context. discourse processes 50. 249–275. carlson, katy, michael walsh dickey, lyn frazier & charles clifton, jr. 2009. information structure expectations in sentence comprehension. the quarterly journal of experimental psychology 62(1). 114–139. demberg, vera & asad sayeed. 2016. the frequency of rapid pupil dilations as a measure of linguistic processing difficulty. plos one 11(1). e0146194. engelhardt, paul e. & fernanda ferreira. 2010. processing coordination ambiguity. language and speech 53(4). 494–509. engelhardt, paul e., fernanda ferreira & elena g. patsenko. 2010. pupillometry reveals processing load during spoken language comprehension. quarterly journal of experimental psychology 63(4). 639–645. fillmore, charles j., paul kay & mary catherine o’connor. 1988. regularity and idiomaticity in grammatical constructions: the case of let alone. language 64. 501–538. frazier, lyn. 1987. syntactic processing: evidence from dutch. natural language & linguistic theory 5(4). 519–559. frazier, lyn. 2018. ellipsis and psycholinguistics. in jeroen van craenenbroeck & tanya temmerman (eds.), the oxford handbook of ellipsis, 253–275. oxford, uk: oxford university press. frazier, lyn & charles clifton, jr. 1998. comprehension of sluiced sentences. language and cognitive processes 13(4). 499–520. harris, jesse a. 2015. structure modulates similarity-based interference in sluicing: an eye tracking study. frontiers in psychology 6. e1839. harris, jesse a. 2016. processing let alone coordination in silent reading. lingua 169. 70–94. harris, jesse a. 2019. alternatives on demand and locality: resolving discourse-linked whphrases in sluiced structures. in charles clifton, jr., janet dean fodor & katy carlson (eds.), grammatical approaches to language processing, 45–75. cham, switzerland: springer. harris, jesse a. to appear. the height of let alone in english: evidence from inversion and contrastive topics. in proceeding of the 40th west coast conference on formal linguistics (wccfl). harris, jesse a. & katy carlson. 2016. keep it local (and final): remnant preferences for ‘let alone’ ellipsis. quarterly journal of experimental psychology 69(7). 1278–1301. harris, jesse a. & katy carlson. 2018. information structure preferences in focus-sensitive ellipsis: how defaults persist. language and speech 61(3). 480–512. harris, jesse a. & sun-ah jun. 2018. using pupillometry to assess prosodic alignment in language comprehension. in sasha calhoun, paola escudero, marija tabain & paul warren (eds.), proceeding of the 19th international congress of phonetic sciences, 2926–2930. harris, jesse a., alexandra lawn & marju kaps. 2019a. investigating sound and structure in concert: a pupillometry study of relative clause attachment. in ashok goel, colleen seifert & christian freksa (eds.), the 41st annual meeting of the cognitive science society, 1880– 1886. harris, jesse a., chie nakamura, bethany sturman & sun-ah jun. 2019b. prosody-meaning proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 126 https://doi.org/10.3765/elm https://www.elm-conference.net/ mismatches in pp ambiguity: incremental processing with pupillometry. poster presented at the 32nd annual cuny conference on human sentence processing. hess, eckhard h. & james m. polt. 1960. pupil size as related to interest value of visual stimuli. science 132. 349–350. hess, eckhard h. & james m. polt. 1964. pupil size in relation to mental activity during simple problem-solving. science 143(3611). 1190–1192. hoof, hanneke van. 2003. the rise in the rise-fall contour: does it evoke a contrastive topic or a contrastive focus? linguistics 41(3). 515–563. hulsey, sarah. 2008. focus sensitive coordination. cambridge, ma: massachusetts institute of technology dissertation. just, marcel a. & patricia a. carpenter. 1993. the intensity dimension of thought: pupillometric indices of sentence processing. canadian journal of experimental psychology 47(2). 310– 339. kahneman, daniel & jackson beatty. 1966. pupil diameter and load on memory. science 154(3756). 1583–1585. kuchinke, melissa l.-h. võ, markus hofmann & arthur m. jacobs. 2007. pupillary responses during lexical decisions vary with word frequency but not emotional valence. international journal of psychophysiology 65. 132–140. kyröläinen, aki-juhani, vincent porretta, jacolien van rij & juhani järvikivi. 2019. pupilpre: tools for preprocessing pupil size data. https://cran.r-project.org/package= pupilpre. laeng, sylvain sirois & gustaf gredebäck. 2012. pupillometry: a window to the preconscious? perspectives on psychological science 7(1). 18–27. lenth, russell v. 2022. emmeans: estimated marginal means, aka least-squares means. https://cran.r-project.org/package=emmeans. r package version 1.7.4-1. mathôt, sebastiaan, jasper fabius, elle van heusden & stefan van der stigchel. 2018. safe and sensible preprocessing and baseline correction of pupil-size data. behavior research methods 50(1). 94–106. scheepers, christoph, sibylle mohr, martin h. fischer & andrew m. roberts. 2013. listening to limericks: a pupillometry investigation of perceivers’ expectancy. plos one 8(9). e74986. schmidtke, jens. 2018. pupillometry in linguistic research: an introduction and review for second language researchers. studies in second language acquisition 40. 529–549. stolterfoht, britta. 2005. processing word order variations and ellipses: the interplay of syntax and information structure during sentence comprehension. leipzig: max planck institute for human cognitive and brain sciences dissertation. toosarvandani, maziar. 2009. letting negative polarity alone for ‘let alone’. in tova friedman & satoshi ito (eds.), proceedings of semantics & linguistics theory (salt), vol. 18, 729–746. amherst, ma. toosarvandani, maziar. 2010. association with foci. university of california, berkeley dissertation. rij, jacolien van, petra hendriks, hedderik van rijn, r. harald baayen & simon n. wood. 2019. analyzing the time course of pupillometric data. trends in hearing 23. 1–22. proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 127 https://doi.org/10.3765/elm https://www.elm-conference.net/ wagner, michael. 2012. contrastive topics decomposed. semantics and pragmatics 5(8). 1–54. winkler, susanne. 2018. ellipsis and prosody. in jeroen van craenenbroeck & tanya temmerman (eds.), the oxford handbook of ellipsis, 357–287. oxford, uk: oxford university press. winn, matthew b., dorothea wendt, thomas koelewijn & stefanie e. kuchinsky. 2018. best practices and advice for using pupillometry to measure listening effort: an introduction for those who want to get started. trends in hearing 22. 1–32. yoshida, masaya. 2018. parsing strategies. in jeroen van craenenbroeck & tanya temmerman (eds.), the oxford handbook of ellipsis, 444–457. oxford, uk: oxford university press. zellin, pannekamp, ulrike toepel & elke van der meer. 2011. in the eye of the listener: pupil dilation elucidates discourse processing. international journal of psychophysiology 81(3). 133–141. proceedings of elm 2: 117-128, 2023 jesse a. harris: the enduring effects of default focus in let alone ellipsis. 128 https://doi.org/10.3765/elm https://www.elm-conference.net/ on the acquisition of either and too naomi francis, shuli jones, leo rosenstein, & martin hackl∗ abstract. this paper presents an experimental investigation of how english-learning children acquire the additive discourse particles either and too. in the target grammar these items exhibit near-complementary distribution conditioned on the polarity of their host sentence. the path leading to that grammar appears to be rather intricate. we present comprehension data showing that for an extended period of time (ages 3–5) learners find both items acceptable in both polarity environments, exhibiting only a weak adult-like tendency of preferring either in negative and too in positive sentences. at age 6 their grammar appears categorical with respect to either in that they no longer tolerate it in positive sentences while still exhibiting only a weak dispreference for too in negative environments. these findings are even more striking in the context of production data. we find that child-directed speech is essentially categorical, providing unambiguous evidence for the adult grammar. moreover, we find essentially categorical, adult-like use of either and too in child production from the earliest stage of development. these observations raise a number of challenges for theories of either and too and for approaches to learning focus particles more generally. perhaps most strikingly, the protracted insensitivity of the learner’s grammar to accumulation of unambiguous evidence constitutes a novel argument from the abundance of evidence for encapsulated learning. keywords. additivity; focus particles; polarity sensitivity; l1 acquisition 1. introduction. this paper presents an experimental investigation of how english-learning children acquire the additive focus particles either and too. just like other non-scalar, additive discourse particles (items such as also, additionally, as well, . . . ), their basic function is to signal that the sentence they attach to is part of a sequence of sentences that constitutes a (potentially partial) answer strategy for addressing a question under discussion (see e.g. beaver & clark 2008). what is special about either and too is that they appear to be designated to mark, respectively, negative and positive sequences. that is, they exhibit (near) complementary distribution, which is conditioned on the polarity of the sentence they accompany. this is exemplified in (1). (1) a. sam is eating cake. sam is eating ice cream too/*either. b. sam isn’t eating cake. sam isn’t eating ice cream either/*?too. one approach to capturing the divergent polarity sensitivity of either and too while maintaining their shared discourse functionality is to assume that they trigger the same canonical additive presupposition (while being assertorically inert), as shown in (2), and to stipulate that they are ∗we are grateful to yadav gowda for his help with building the scripts that extracted and coded our childes data, to anya keomurjian, abena peasah, and mika thakkar for constructing stimuli for the adult experiment, and to steve worthington for stats guidance. this research was supported by the social sciences and humanities research council of canada (doctoral fellowship to the first author). authors: naomi francis, university of oslo (n.c.francis@iln.uio.no), shuli jones, massachusetts institute of technology, leo rosenstein, massachusetts institute of technology, & martin hackl, massachusetts institute of technology. proceedings of elm 1: 159-171, 2021 c©2021 naomi francis, shuli jones, leo rosenstein, and martin hackl published by the lsa with permission of the author(s) under a cc by license. 159 https://doi.org/10.3765/elm https://www.elm-conference.net/ (morpho-) syntactically specified for specific environments. either would be marked to occur with negative sentences, taking necessarily wide scope over the negative operator, while too would be confined to positive sentences either via designated marking or as the result of competition with its more marked counterpart either.1 (2) jeither/took(p) = λw: ∃q ∈ alt(p) s.t. q(w)=1. p(w)=1 rullmann (2003) provides compelling arguments against such a treatment, however, and instead proposes an analysis on which either and too are not synonymous. concretely, while too receives a treatment like any run-of-the-mill additive particle, either is assumed to trigger an “anti-additive” presupposition demanding that there is an alternative to its host clause that is false, (3).2 (3) jeitherk(p) = λw: ∃q ∈ alt(p) s.t. q(w)=0. p(w)=1 (3) by itself does not prevent either from being attached to positive sentences, of course. to make that impossible rullmann (2003) needs to assume, moreover, that either is marked as an npi, hence confined to occur in the scope of a suitable negative operator.3 this ensures that the host clause of either is negative. moreover, since either presupposes that there is an alternative (antecedent) proposition to p that is false results in either being able to serve as a signal for a sequence of negated propositions forming an answer strategy. the fact that too can mark a positive sequence follows directly. to account for the fact the too cannot mark a negated sequence requires, however, a further assumption. rullmann (2003) proposes that too is preferentially attached low in the structure, lower, importantly, than the attachment site of sentence negation. attaching too below negation in (1-b) triggers a positive presupposition which the context does not satisfy, however, giving rise to infelicity.4 by the same token, this proposal predicts that speakers might find (1-b) with too acceptable to the extent to which they are able to access a parse with high attachment of too. in that case the host clause happens to express a negated proposition which is licit as long as the context provides a true focus alternative to it. since this is the case in (1-b) wide scope attachment of too would produce a felicitous sequence, offering a perspective on why the unacceptability in (1-b) is typically judged less severe than the unacceptability of (1-a). although the proposal in rullmann (2003) is not without problems and there are competing theories in the literature, notably the alternatives-based analysis of ahn (2015), we think that it provides a useful framework within which to explore questions about how either and too might be acquired. in particular, we are interested in finding out whether learners initially adopt the most basic additive semantics for either as in (2). evidence for such a stage in development 1this could be implemented via morphological suppletion which is conditioned on the syntactic position of the additive particle, as in könig (1991), or via some form of feature probing mechanism. 2how exactly alternative propositions are generated is an open question in the literature. for our purpose it will be sufficient to follow beaver & clark (2008) and assume that the the focus marking inside the host clause reflects which qud the sequence is meant to address. 3see rullmann (2003) for details. since the question of which negative operators can license either is not central to our purpose, we will organize our discussion around sentence negation and abstract away from any complications that arise when considering other licensors. 4however, in case there is a positive alternative provided by the context as in (i) too is predicted to be licensed. (i) mary ate the lasagna, but she couldn’t eat the spaghetti, too. (rullmann 2003; 389) proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 160 https://doi.org/10.3765/elm https://www.elm-conference.net/ would come in the form of either being fully interchangeable with too across positive and negative environments. if such a stage can be identified, we would then hope to trace the time course and the evidence it takes to learn their near-complementary distribution. of particular interest is whether learners go through a stage predicted by the proposal that rullmann (2003) ends up rejecting. if so, these learners should expect either and too to exhibit complementary distribution in basic sentences such as (1) but they would not yet know that either is, in fact, an npi. to investigate these questions we first conducted a comprehension experiment employing a comparative felicity judgment task. this task is well-suited for probing directly whether speakers prefer either over too in positive environments and too over either in negative environments. we find that children do not behave in accordance with rullmann (2003) or the simpler proposal rejected therein. indeed, our findings are puzzling under all theories of either and too we are aware of. we then conducted a corpus study on child-directed speech to see whether their comprehension could be understood in terms of the data available to the learner being noisy, uninformative or even misleading. interestingly, our corpus study revealed quite the opposite: child-directed speech is in effect categorical with respect to the polarity-dependent distribution of either and too. finally, we conducted a corpus study on child speech to examine whether their productions of either and too reflect their comprehension system. we find, however, that child speech is in fact adult-like even at the earliest stage in development. this puts our comprehension data in even starker relief. 2. child experiment. 2.1. method. to assess children’s understanding of the polarity-conditioned distribution of either and too, we employed a comparative felicity judgment task (chierchia et al. 2001, foppolo et al. 2012). this task presents participants with a forced choice between either and too in a given environment. for positive sentences, we predict that adults will always choose too over either, since there is no way for either to be licensed there. for negative sentences, a prediction based on rullmann (2003) is more nuanced; speakers who cannot access a wide-scope lf for too should have a categorical preference for either over too. for speakers who can access a wide-scope lf for too, the rate of selecting too will depend on them judging the either competitor to be nevertheless better. if they are roughly comparable, we might expect non-categorical selection of both items; if they are significantly different, as rullmann suggests, we might expect a fully categorical pattern.5 in each trial, we used a laptop to present a scene like those exemplified in figure 1 to the participant. one experimenter, playing the role of narrator, introduced the scene and provided an appropriate antecedent for the additive particle. then two puppets provided competing sentences completing the description; the puppets’ utterances were identical except for the additive item used (either vs. too). children were then asked to indicate which puppet “said it better”. the experiment consisted of eight target items and four filler items, balanced for polarity. positive and negative versions of each target item were created. for each target item, half of the participants saw the positive version and half saw the negative version; the fillers were seen by all. the ordering of positive and negative trials, the order in which the puppets spoke, and the additive particle used by the first speaker were pseudo-randomized to avoid unwanted patterns. we recruited 57 participants aged 3–7 years in the boston area at the museum of science and 5since this is an empirical question, we ran an adult control study reported in section 3. proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 161 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: examples of positive (l) and negative (r) target items local daycares. of these, 47 l1 english acquirers completed the experiment: nine 3-year-olds, fourteen 4-year-olds, fifteen 5-year-olds, eight 6-year-olds, and one 7-year-old. the 7-year-old was excluded from our analysis, leaving data from 46 participants. a parent or guardian provided informed consent for each child, and each child received a sticker in thanks for their participation. 2.2. results. the mean rates at which children selected each item in each environment are shown in figure 2. figure 2: mean rate of either/too selection, by polarity environment and age to examine the distribution of our responses we constructed a linear mixed effects logit model in r (r core team 2019) using the lme4 package (bates et al. 2015) with fixed effects of polarity (pos vs. neg) and age (3–5 vs. 6). this model showed main effects of polarity and age; children were more likely to select either in negative sentences than in positive sentences (pr(>|z|) = 0.000171 ***), and 6-year-olds were less likely to select either overall (pr(>|z|) = 0.009378 **). the model also revealed a significant interaction between polarity and age (pr(>|z|) = 0.036657 *); the contribution of this interaction was confirmed with a likelihood ratio test (p < 0.05 *). proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 162 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.3. discussion. we take the comparative felicity judgment task to reflect comprehension, as it tests participants’ responses to utterances of which they are not the author. the results reported above show that children are not adult-like in their treatment of either and too. from the ages of 3 to 5, children exhibit a non-categorical response pattern, accepting both items in both environments. however, their behavior is not random; they select either more often than too in negative environments and too more often than either in positive environments. that is, they trend toward the adult grammar. importantly, this trending pattern does not reflect a population-level split, with some children being categorical in one direction and others being categorical in the other; the non-categorical trending pattern is attested within subjects. that this behavior is non-adult-like is transparent for either, but the interpretation of the too data is more involved. informal fieldwork with native english speakers and an adult felicity rating study reported in section 3 suggest that too in negative sentences is judged to be quite infelicitous. a small (n=4) replication of the above experiment with adults confirmed that they deliver categorical judgments on this task. moreover, some children (n=8) explicitly indicated that both items were equally good in the sentences they were presented with, answering “both” or “neither” to the experimenter’s question “who said it better?”. we take this to indicate that children’s treatment of too in our task is non-adult-like. the non-categorical trending behavior continues virtually unchanged until the age of 6, when children suddenly stop selecting either in positive sentences. this is presumably responsible for the statistical effect of age, which indicated that 6-year-olds are less likely to accept either overall. the patterns of behavior exhibited by the 3–5-year-olds on the one hand and the 6-year-olds on the other both raise interesting questions. what knowledge state does the 3–5-year-olds’ behavior reflect? their different response patterns in positive and negative environments suggest they are not ignorant of the role of polarity in conditioning these items’ distribution, but neither do they faithfully reproduce its categorical nature. the stability of this trending behavior is also striking; the pattern of responses displayed by the 3-year-olds in our sample is virtually identical to that displayed by the 5-year-olds, despite the fact that the latter group has had two more years’ worth of exposure to the distribution of these items in the input. this suggests that whatever is responsible for the change in behavior that occurs between ages 5 and 6 is not merely the result of the steady accumulation of evidence; this constitutes a novel argument for encapsulated learning. how should we characterize the learning pattern displayed here? the early grammar of additivity seems to have compatibility of either and too as one of its core properties. at first glance, this might suggest that children initially hypothesize that these items share a basic additive semantics (i.e., (2)). however, on such a view it is difficult to explain the trending behavior, and equally difficult to explain its stability. why it is so difficult for children to master the restricted distribution of these items? this fact becomes even more striking once we understand the input is robustly categorical (see section 4). from the perspective of rullmann’s (2003) account, we would ask what it takes to move from a simple additive semantics for either (requiring a true antecedent, as in (2)) to the non-standard one (requiring a false antecedent, as in (3)).6 the answer offered by this framework appears simple enough at first blush: children must realize that either is an npi. what is left open, however, is why a wide-scope theory of either is not entertained at all, not even at an 6if we take our data seriously, it appears that the children we tested never entertain a simple könig-style approach. proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 163 https://doi.org/10.3765/elm https://www.elm-conference.net/ intermediate state for which it is plausible to assume that learners have not yet encountered data that decisively argue for the npi-treatment of either. 3. adult control. 3.1. method. to precisify what the target grammar of too allows, we ran a variation of the experiment described in the previous section on an adult population. our goal was to determine whether too is acceptable in negative sentences with negative antecedents on the adult grammar. a direct replication of the child experiment may be too blunt a tool to investigate this question, as it does not distinguish between what is dispreferred because it is infelicitous and what is dispreferred because it is merely less felicitous than its competitor. to give participants an opportunity to indicate a difference between what is allowed but dispreferred by their grammar vs. what is not allowed, we conducted a felicity rating experiment. our adult study was conducted online. participants were shown scenes similar to those used in the child experiment and asked to rate the naturalness of positive and negative sentences with either and too (each following an appropriate antecedent) on a likert scale with values ranging from 1 (“extremely unnatural”) to 7 (“perfectly natural”). the experiment consisted of 32 latin-squared target items and 40 filler items, both balanced for polarity. in each trial the scene and antecedent sentence appeared on the screen; to ensure that participants were actively engaged in the task and had time to read the text, they had to click their mouse or press a key after a 2000 ms embargo period to advance to the target sentence and then to the rating scale. participants were rejected if they i) failed to complete their rating within 7000 ms of the scale appearing in more than 10% of trials, or ii) rated more than two of eight ungrammatical benchmarking fillers above a 2. we used ibex farm (drummond 2018) and penncontroller (zehr & schwarz 2018) with amazon’s mechanical turk to recruit participants and deploy the experiment. we collected data from 51 l1 u.s. english-speaking adults. of these, 6 were excluded due to their performance on the benchmarking fillers, leaving data from 45 individuals. participants gave informed consent and received 3.85 usd. 3.2. results. to account for possible variation in how participants used the 7-point scale, we normalized the ratings using z-scores. the resulting mean ratings are plotted in figure 3. a maximally-specified convergent linear mixed effects model of z-scored ratings constructed in r using package lme4, with fixed effects of polarity (pos vs. neg) and particle (either vs. too), reveals a significant polarity by particle interaction (p < 0.001 ***). either was judged to be more acceptable in negative than in positive sentences; the opposite was true for too (t = 76.00). 3.3. discussion. we found that the ratings depended on a highly significant interaction between the polarity of a sentence and the additive item used. it appears that ratings for negative too sentences cluster with those for positive either sentences rather than the perfectly acceptable positive too and negative either sentences; note that this matches rullmann’s (2003) description of the data rather well. these results suggest that adults indeed treat too as degraded in negative environments, albeit less so than either in positive environments. we take this to indicate that even at age 6 children are not adult-like in their treatment of too in negative environments. proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 164 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 3: z-scores of ratings 4. corpus study: child-directed adult speech. we have observed that children’s comprehension of either and too is non-adult-like until at least age 6. it is important to know whether this protracted development could be attributed to the input; do children not receive sufficient clear evidence about the distributions of either and too, or do they have difficulty using the evidence? to determine whether children’s non-categorical trending behavior can be attributed to properties of the input, we conducted a corpus study of child-directed speech. using custom r scripts, we extracted every instance of additive too and either from childes corpora of typically-developing, north american english-acquiring children spoken by mother, father, sister, brother, aunt, uncle, grandmother, grandfather, family friend, teacher, or adult (macwhinney 2000). each token was coded for the polarity of the environment in which it appeared. the results are given in the following table 1. children hear approximately ten times item total positive negative unclear either 701 28 (3.99%) 670 (95.58%) 3 (0.43%) too 7896 7782 (98.56%) 103 (1.30%) 11 (0.14%) table 1: child-directed utterances of additive either/too in childes more occurrences of additive too than either. the input is overwhelmingly categorical, with very few instances of unlicensed either and too.7 it should be noted, however, that if children ask themselves not where to insert either or too but rather which item to insert in a positive or negative environment, the data available to them in negative environments is somewhat noisier than that 7the picture is in fact even clearer than these numbers suggest; 23 of the 28 positive either tokens were found in sentences like me either!, uttered in response to a negative antecedent, while 42 of the negative too tokens were sentences like don’t i get some too?, following a positive antecedent. proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 165 https://doi.org/10.3765/elm https://www.elm-conference.net/ available in positive ones.8 even so, the input should provide very few opportunities for confusion about the distribution of either and too. 5. corpus study: child speech. in the previous section, we observed clear categorical behavior in child-directed speech that contrasts with the children’s non-categorical behavior profile in the comparative felicity judgment task. this raises a question: what does the production system look like? does it mimic the input or does it reflect the comprehension system? we repeated the procedure described above for the target children in each of the corpora investigated. the results are given in table 2. children’s productions, like adults’ productions, item total positive negative unclear either 242 22 (9.09%) 210 (86.78%) 10 (4.13%) too 5059 5007 (89.95%) 42 (0.83%) 10 (0.20%) table 2: child utterances of additive either/too in childes are categorical; although children were willing to accept the puppets’ productions of too in negative environments and (for the younger group) either in positive environments, they themselves very rarely use these items in environments where they are not licensed.9 children’s productions thus broadly mirror the adult productions that they are exposed to, although the former are somewhat noisier. it is worth noting that, although children (like adults) produce more utterances containing additive too than either, both particles appear to follow the same growth curve, figure 4. figure 4: children’s cumulative productions of either and too 6. discussion. our investigation has revealed a number of unexpected and indeed rather noteworthy features of children’s developing understanding of either and too. the most significant ones are: i) from ages 3–5, children display a stable, non-categorical preference in the direction of the adult grammar, ii) between ages 5 and 6, there is an abrupt transition to adult-like categorical behavior in one half of the paradigm (either’s restriction to negative environments) but not the other, and iii) there is an asymmetry between children’s comprehension and production, with the former 8the percentages of adult-grammar consistent data would be 86.68% for negative environments and 99.74% for positive ones. we are grateful to jesse snedeker for pointing this out to us. 9as with the adult production data, most of the positive either tokens (18 out of 22) were sentences like me either!, uttered in response to a negative antecedent. proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 166 https://doi.org/10.3765/elm https://www.elm-conference.net/ lagging behind the latter. we believe that these findings present serious challenges for all existing theories of (adult) either and too. in keeping with our previous discussion, we will explicate the nature of these challenges against the backdrop of rullmann (2003), attempting along the way to distill more general desiderata for any theory which hopes to account not only for the adult pattern but aims to provide a principled basis for explaining how these items are acquired. from the perspective of the target grammar advocated in rullmann (2003), we expect the acquisition path for either and too to unfold roughly along the following lines. upon first encounter of either and too learners should entertain a run-of-the-mill additive semantics as in (2) for both items. this is because the treatment in (2) is arguably the simplest way of specifying an additive particle and it is independently needed to account for particles like also.10 the resulting grammar predicts either and too to be completely interchangeable across all environments and so cannot account for their much more restricted distribution. next, as learners begin to record the distributional peculiarity of either and too in their input data, which, recall, we found to be essentially categorical, they should amend the respective entries in such a way that either is confined to negative environments and too to positive. the adjustment to the grammar might take the form of specifying each item morpho-syntactically for a suppletion rule along the lines of könig (1991), where either is simply the spell-out of the particle when it takes scope over a relevant negative operator. alternatively, either might be equipped with a syntactic selection feature that can only be discharged if the host clause contains a negative operator in the scope of either. specifying a lexical item morpho-syntactically for a restricted set of environments is by no means unprecedented in the grammar and so should be a readily available option for the learner. moreover, since the required morpho-syntactic amendment can take place without any change to the basic additive semantics of our particles we expect a stage in development that is well-characterized by a wide-scope either grammar. it would exhibit a fully complementary distribution of our particles for simple sentences like those in (1) and remain in place until the learner encounters more complex sentences that cannot be captured by such a grammar. one type of data that would provide compelling evidence against the wide-scope either grammar is exemplified in (4), where the negative operator is in a higher clause but the presupposition triggered by either does not contain the matrix verb. (4) john refused to go to church. his parents were so angry that they did not permit him to go to the soccer game either. (rullmann 2003; 353) the additive inference in (4) is quite transparently satisfied by the proposition that john didn’t go to church. on the wide-scope either theory, however, (4) is incorrectly predicted to trigger the presupposition that john’s parents did not permit him to go to church. this is because on the wide-scope theory, either necessarily outscopes the negative operator of its host sentence. only in that position is it possible for either to generate the correct additive inference in simple cases such as (1). for more complex cases like (4) this assumption fails. to account for the correct presupposition the learner needs to assume that either, in fact, takes narrow scope with respect to negation, which in turn requires abandoning the basic additive semantics for either in favor of its 10note furthermore that while the proposal in rullmann (2003) implies a hypothesis space that allows for an “antiadditive” semantics for additive discourse particles, one where the presupposition requires a false antecedent, there are no negative-positive sequences in the data which would unambiguously call for such an entry. proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 167 https://doi.org/10.3765/elm https://www.elm-conference.net/ anti-additive counterpart in (3). it also requires marking either morpho-syntactically for npi-status to ensure that it cannot occur with positive hosts. this adjustment to the grammar is clearly more involved than the previous one. moreover, since the type of evidence that would trigger the change involves more complex constructions which are presumably rarer in the input, it is expected under such an account that the switch to the npi-grammar of either could occur quite late in development. our findings are only partially consistent with these expectations. to begin with, we do not see a stage in development where both items are completely interchangeable across positive and negative sentences. of course, it is possible that such a stage occurs earlier, before the youngest age at which we were able to gather production and comprehension data. looking at the data we do have, we see that even our youngest participants exhibit behavior that is not fully compatible with this prediction. this is most clearly documented by our production data where we find that even the earliest productions of either and too are adult-like in their polarity sensitivity. turning to our comprehension data, the picture gets more nuanced and more interesting. recall that we find that 3–5-year-olds judged both particles to be contextually appropriate with both sentence types. by itself, this would be expected if they hypothesized the basic additive semantics in (2) for both particles. however, what is not expected is that these children also exhibit a weak adult-like trend towards preferring either with negative hosts and too with positive hosts, and that they do so essentially unchanged for an extended period of time (three years of development). both aspects of this development are quite puzzling from the perspective of rullmann (2003) and even more so from the perspective of the simpler grammar in könig (1991). the fact that these children exhibit some sensitivity as evidenced by their trending behavior shows that they have caught on to the fact that polarity matters somehow for the distribution of either and too. even more striking, we know from their productions that they view the effect of polarity on the choice of the particle for the purpose of production as decisive. why then are they not able to execute a rather simple adjustment to the lexical entries of these items that would result in a grammar that matches their input as well as their own productions? put differently, we have strong indications form our learner’s production and comprehension that the polarity of the host sentence matters for the status of either and too. nevertheless, our learners seem to view them as compatible with both environments in comprehension, attributing to them the ability to trigger a contextually suitable additive inference irrespective of the polarity of the host sentence. this is of course starkly non-adult-like. and it is entirely mysterious from the perspective of the wide-scope grammar of either advocated e.g. in könig (1991) but also from the perspective of the npi-grammar proposed in rullmann (2003) since that theory offers no reasons why learners shouldn’t entertain a wide-scope grammar at an intermediate stage of acquisition when they may have taken note of the polarity effect on the distribution of either and too in basic sentences but have not yet encountered data such as (4) that are transparently incompatible with a wide-scope grammar of either. arguably the most striking aspect of this pattern is that it persists over three years of development in the face of categorical and readily available data to the contrary. for some reason, learners seem to not make use of this evidence; the only reason for this seems to be that the wide-scope either grammar is never entertained by learners to begin with. that is, it is not part of the hypothesis space for additive discourse particles.11 if this diagnosis is correct, we have here a novel case of 11this would predict that there are no such items in the world’s languages. to our knowledge, this has not been proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 168 https://doi.org/10.3765/elm https://www.elm-conference.net/ encapsulated learning, one that is based on the abundance rather than the poverty of the stimulus.12 next, we turn to the change that occurs between ages 5 and 6, which affects either in that it is now excluded from positive environments but does not seem to affect too, which is still readily tolerated in negative environments. rullmann (2003) offers a perspective for the timing of this change since it is quite plausible that informative data such as (4) are rare in the input and, furthermore, require a well-developed grammar to be fully appreciated by a learner.13 it is less clear, however, whether 6-year-olds judging too in negative sentences to be fully felicitous (albeit at a slightly lower rate than either) fits as well. on rullmann (2003) too is disallowed in negative sentences because the required high-scope lf is only marginally accessible due to a parsing preference to attach sentence final particles low and blocking by either which can convey the same presupposition without violating the parsing preference. note that our adult control data bear out this prediction quite nicely. the behavior of our 6 years olds in the comparative felicity judgment task seems to indicate, however, that too does not compete with either in negative sentences even though either is already understood to be an npi at that age. the only way this could be explained within rullmann (2003) is to assume that 6-year-olds have not yet adopted the low-attachment preference for sentence-final particles.14 the non-categorical (non-adult-like) behavior observed in the 3–5-year-olds’ comprehension performance contrasts sharply with their fully categorical (adult-like) productions. while asymmetries between comprehension and production in language acquisition are not uncommon, such asymmetries usually go in the opposite direction, with comprehension maturing first. one attractive way of deriving the appearance of comprehension lagging behind production is to exploit the speaker’s advantage of knowing the message;15 however, it is not clear that our case is amenable to such an explanation, as the stimuli in the comprehension experiment should have left no doubt about the intended message and the context provides no invitation to infer an enriched meaning. we suspect that our asymmetry is instead a reflection of formal complexity.16 if two structures can express the same meaning, with one being more complex than the other, we might expect that children would systematically choose to produce the simpler structure without necessarily insisting on systematically investigated. 12see e.g. babyonyshev et al. (2001) for a similar case. 13our own corpus work on child-directed speech is not fine-grained enough to examine this prediction and we will have to leave this issue for future research. 14whether this is plausible has, to our knowledge, not been investigated in literature. 15see hendriks (2014) for discussion of a variety of different relevant cases and an optimality theory-based perspective under which the advantage for production is a consequence of a non-adult ranking of grammatical constraints. 16we do not present a detailed comparison with ahn’s (2015) account here, but we do not believe that it has a significant advantage in capturing our data. in a nutshell, ahn proposes that either and too assert disjunction and conjunction, respectively, between the prejacent (p) and a covert propositional anaphor (q) that is presupposed to be a focus alternative of the prejacent. she derives the npi behavior of either via exhaustification of obligatorily-activated subdomain alternatives in the style of chierchia (2006). an alternatives-based approach along these lines might seem to provide an explanation for the children’s behavior, as children under the age of 6 have been found to be non-adultlike in their ability to access certain alternatives; indeed, it has been claimed that this restricted set of alternatives allows at least some children in this age group to systematically strengthen overt disjunction to a meaning that is fully equivalent to conjunction, though the structure that produces this meaning is more complex (singh et al. 2016). however, our experimental setup is known to circumvent this kind of problem, with children systematically selecting the more informative of two options presented to them (chierchia et al. 2001, foppolo et al. 2012). proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 169 https://doi.org/10.3765/elm https://www.elm-conference.net/ the same degree of economy from other speakers. for example, suppose that children’s developmental path takes them through a stage where a sentence can have the same meaning regardless of which additive item is inserted, but the either-sentence has a more complex lf than too-sentence if it is positive and the reverse holds in negative environments. if children always produce the simplest structure that will convey their intended message, we should expect them to systematically produce too in positive sentences and too in negative ones. importantly, this leaves open the possibility that children will accept other speakers’ productions of both items in both environments, as we observed in the comprehension task, without necessarily inferring an enriched meaning in case the speaker chooses the less economical option. references ahn, dorothy. 2015. the semantics of additive either. in eva csipak & hedde zeijlstra (eds.), proceedings of sinn und bedeutung 19, 20–35. babyonyshev, maria, jennifer ganger, david pesetsky & kenneth wexler. 2001. the maturation of grammatical principles: evidence from russian unaccusatives. linguistic inquiry 32(1). 1–44. https://doi.org/10.1162/002438901554577. bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67(1). 1–48. https://doi.org/10.18637/jss.v067.i01. beaver, david i. & brady z. clark. 2008. sense and sensitivity: how focus determines meaning explorations in semantics. oxford, uk: wiley-blackwell. chierchia, gennaro. 2006. broaden your views: implications of domain widening and the “logicality” of language. linguistic inquiry 37(4). 535–590. https://doi.org/10.1162/ling.2006.37. 4.535. chierchia, gennaro, stephen crain, maria teresa guasti, andrea gualmini & luisa meroni. 2001. the acquisition of disjunction: evidence for a grammatical view of scalar implicatures. in bucld 25 proceedings, 157–168. somerville, ma: cascadilla press. drummond, alex. 2018. ibex farm. web. https://spellout.net/ibexfarm/. foppolo, francesca, maria teresa guasti & gennaro chierchia. 2012. scalar implicatures in child language: give children a chance. language learning and development 8. 365–394. https://doi.org/10.1080/15475441.2011.626386. hendriks, petra. 2014. asymmetries between language production and comprehension. dordrecht: springer. könig, ekkehard. 1991. the meaning of focus particles: a comparative perspective. london and new york: routledge. macwhinney, brian. 2000. the childes project: tools for analyzing talk. mahwah, nj: lawrence erlbaum associates 3rd edn. r core team. 2019. r: a language and environment for statistical computing. r foundation for statistical computing vienna, austria. https://www.r-project.org/. rullmann, hotze. 2003. additive particles and polarity. journal of semantics 20. 329–401. https://doi.org/10.1093/jos/20.4.329. singh, raj, ken wexler, andrea astle-rahim, deepthi kamawar & danny fox. 2016. children proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 170 https://doi.org/10.3765/elm https://www.elm-conference.net/ interpret disjunction as conjunction: consequences for the theories of implicature and child development. natural language semantics 24(4). 305–352. https://doi.org/10.1007/s11050016-9126-3. zehr, jérémy & florian schwarz. 2018. penncontroller for internet based experiments (ibex). web. https://doi.org/10.17605/osf.io/md832. proceedings of elm 1: 159-171, 2021 naomi francis, shuli jones, leo rosenstein, and martin hackl: on the acquisition of either and too. 171 https://doi.org/10.3765/elm https://www.elm-conference.net/ superlative-modified numerals and negation: a multiply negotiable cost teodora mihoc & kathryn davidson* abstract. comparative-modified numerals (cmns) and superlative-modified numerals (smns) are reported to exhibit a polarity sensitivity contrast: unlike cmns, the use of smns is said to be sensitive to embedding under negation. this contrast is however neither well studied nor well understood, such that the existing views in the literature disagree vastly on both the basic facts and related expectations. in this paper we investigate this contrast in three offline experiments. we show that there is strong empirical support for this reported contrast under negation, but none of the existing analyses of the contrast can capture it in full, though we do seem to require insights from each. keywords. comparative-modified numerals, superlative-modified numerals; negation 1. introduction. 1.1. the polarity sensitivity contrast. naively speaking, given reference to the same scale, comparative-modified numerals (cmns; e.g., more/less than 3) and superlative-modified numerals (smns; e.g., at least/most 3) are pairwise truth-conditionally equivalent (more than 2 = at least 3 = 3 or 4 or . . . ). however, they are reported to differ in multiple ways. one reported contrast has to do with polarity sensitivity. this contrast is usually described as follows: cmns are fine in all downward-entailing (de) environments, but smns have a mixed profile: their acceptability is degraded in some de environments, for example a negative declarative (decl-neg), but acceptable in other de environments, for example, the positive antecedent of a conditional (antcond-pos) or restriction of a universal (restuniv-pos) (geurts & nouwen 2007, a.o.). (1) jo didn’t call 3more than 2 / # at least 3 people. (not > 3cmn / # smn) (2) if jo called 3more than 2 / 3at least 3 people, she passed. (if > 3cmn / 3smn) (3) everyone who called 3more than 2 / 3at least 3 people passed. (every > 3cmn / 3smn) 1.2. four theoretical positions. the existing literature implicitly or explicitly espouses the following four views with respect to the polarity sensitivity contrast. except for the first one, which denies the contrast, each of these views has a story for the contrast in decl-neg and the non-contrast in antcond/restuniv-pos, as well as predictions for a negative antecedent of a conditional / restriction of a universal (antcond/restuniv-neg: if jo didn’t call at least 3 people, she passed / everyone who didn’t call at least 3 people passed) and at least one other prediction. we briefly review each below. v0: no contrast. this (admittedly straw-man) view holds that there is no contrast. v1: processing cost. this view is as follows: smns are degraded in decl-neg and fine in antcond/restuniv-pos because smns incur an extra processing cost (see, e.g., alexan*we would like to thank athulya aravind, gennaro chierchia, andreea nicolae, and steven worthington (the institute of quantitative social sciences at harvard), as well as members of the experimental syntax & semantics lab / language acquisition lab at mit and of the meaning and modality lab and/or of the language and cognition lab at harvard, and audiences at xprag 2017 and elm 1. any errors are our own. authors: teodora mihoc, harvard university (tmihoc@fas.harvard.edu) & kathryn davidson, harvard university (kathryndavidson@fas.harvard.edu). proceedings of elm 1: 212-223, 2021 c©2021 teodora mihoc and kathryn davidson published by the lsa with permission of the author(s) under a cc by license. 212 https://doi.org/10.3765/elm https://www.elm-conference.net/ dropoulou 2018 and refs. therein) and the first environment contains negation, which adds to this cost (wason 1961, clark & chase 1972), but the other two environments do not. this view predicts that smns in antcond/restuniv-neg should be degraded also, perhaps even more, as these environments not only contain negation but are also even more complex. this view also suggests that the de modifiers might also generally be worse, as they are also in some sense negative. v2: monotonicity and evaluativity (cohen & krifka 2014). this view goes as follows: smns are degraded in decl-neg and fine in antcond/restuniv-pos because they have two lexical meanings, one sensitive to monotonicity, which requires that the smn be in an upward-entailing (ue) environment, a condition violated in both declneg and antcond/restuniv-pos above, and one sensitive to evaluativity, which requires that the property that the smn combines with be pragmatically positive, a condition met in antcond/restuniv-pos above, where the positive continuation passed forces the otherwise neutral call people to be understood as positive. this view predicts that, on the meaning sensitive to monotonicity, smns in antcond/restuniv-neg should always be fine, as these doubly de environments are altogether ue. this view also predicts that, on the meaning sensitive to evaluativity, an smn in either antcond/restuniv-pos or antcond/restuniv-neg may be degraded: if you click at least twice, # the system will crash / if you don’t click at least twice, # you will get a discount is degraded because the negative / positive continuation forces the antecedent to be understood as negative / positive, which forces the predicate click to be understood as negative. v3: monotonicity alone (spector 2015, mihoc 2020). this view goes as follows: smns are degraded in decl-neg and fine in antcond/restuniv-pos because smns are sensitive to downward monotonicity and the first environment is de at all levels whereas the other two actually contain an ue presupposition (e.g., if jo called. . . / everyone who called. . . presuppose that it is possible that jo/someone called). this view predicts that smns should be fine in antcond/restuniv-neg, as these environments are ue. this view also predicts that smns should be fine, for example, under a negated factive (e.g., tim doesn’t know that jo called . . . ), as this also yields a de environment with an ue presupposition (that jo called . . . ), or under two negations (e.g., tim doesn’t know that jo didn’t call . . . ), as this also yields an ue environment. 1.3. goal of this paper and plan. the four views above are very different. in this paper we report on three offline experiments in which we tested their basic assumptions and expectations. in exp. 1 we checked the basic patterns, trying to decide between v0 and v1-3 and between v1 and v2-3 (§2). in exp. 2, we tested a prediction from v2 (§3). and in exp. 3 we tested a prediction from v3 (§4). as we show, our findings reject v0 and support a mixture of v1-3 (§5). 1.4. general notes on methods and results. methods. the sentences we wanted to test pose a number of specific challenges: (1) these sentences are complex, with multiple logical operators, and inherently awkward. (2) at least in a positive declarative, the epistemic state of the speaker might introduce a known confound (smns require an ignorant speaker, though cmns are fine with either knowledge or ignorance; nouwen et al. 2019 and refs. therein). (3) in a negative declarative the target narrow scope, how many?, reading can be avoided by interpreting the smn with a wide scope, specific reading (mayr 2013). to address all these points, we decided to adopt a card game scenario inspired by cremers & chemla (2017), making it clear that the characters in the game are always talking about how many? proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 213 https://doi.org/10.3765/elm https://www.elm-conference.net/ cards, and never about specific cards, and keeping the epistemic state of the speaker transparent and constant. aside from these, an additional challenge is as follows: (4) the nature of the polarity sensitivity contrast is not clear—these sentences appear to be syntactically well-formed and, with some effort, one can compute their meaning. to address this, we decided to ask participants to provide what we term comprehensibility judgments: do you think x will understand what y said? all of (1)-(4) are enforced via pictures in every trial. in addition to these, there were further difficulties, namely: (5) varying the numeral might add further complexity. to address this we decided to keep it constant (always three). finally, (6) the contrast is fairly subtle. to maximize chances of detecting it, since we had no reason to be concerned about conscious access to the task affecting responses, we chose not to use any fillers, and merely to present the stimuli in a new, random order each time. concrete examples of the tasks and stimuli are given in what follows, and printouts of the full surveys online at https://osf.io/6gpu3/. finally: (7) we wanted a diverse population. to address this, we recruited participants from amazon mechanical turk (self-reported native speakers of english, different for each of exp. 1-3, and paid $2, $1, and $1), and provided a link to the survey, presented via qualtrics software. results. the results were analyzed in r. in each case we first report the raw results in the form of a plot, then also the result of fitting logistic mixed effects regression models, commenting in particular on general effects, contrasts between the two modifier types, and contrasts within the modifier types. our data and analysis scripts are available at https://osf.io/6gpu3/. 2. experiment 1. 2.1. goal. to test the basic patterns for decl-neg and antcond/restuniv-pos, as well as for antcond/restuniv-neg, as a way to tease apart v0 from v1-3 and v1 from v2-3. 2.2. methods. task: see figure 1a. sample trial: see figure 1b. trial summary: there were 24 trials, obtained by crossing the following factors: env = embedding environment (decl = declarative; antcond = antecedent of conditional; restuniv = restriction of universal); pol = polarity of env (pos = positive, neg = negative); and modtype = modifier type (comp = comparative, sup = superlative) x modmon = modifier monotonicity (ue, de), which yield mod = modifier (morethan, lessthan, atleast, atmost). see figure 1c. participants: 99, of which 3 were excluded (due to unrealistic end time, or providing the same answer in all trials). 2.3. results. for the raw means see figure 1d and for model results see tables 1a-1c. general effects: for all environment types, and for both cmns and smns, ratings were significantly lower for modmon=de, pol=neg, the interaction of modmon=de with pol=neg, and also (marginally) for env=restuniv. see table 1a. contrasts between modifier types: for the same level of monotonicity, smns in decl-pos were the same as cmns; in decl-neg were much worse than cmns; in a antcond/restuniv-pos were the same as cmns, except for at most in a restuniv-pos, which was worse; and in antcont/restuniv-neg were slightly worse than cmns. see table 1b. contrasts within modifier types: cmns degraded from declneg to antcond/restuniv-neg, but smns did not. see table 1c. 2.4. discussion. smns were much worse than cmns in decl-neg but on a par with cmns in antcond/restuniv-pos. this argues strongly against v0 and in favor of v1-3. smns were in fact significantly worse than cmns in every neg condition. this might seem to support v1 proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 214 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) exp. 1 task. in this survey you will answer questions about a group of friends playing a game. at the beginning of the game each player gets dealt a hand of seven cards. after taking a quick look at them, they must place the cards face down and try to remember their hands. then they take turns giving clues about their hands to the other players in the form of statements describing their hands. you will see what a player remembers about his/her cards and the statement s/he makes, then you will be asked if you think the other players will understand what s/he said. note: a or a means that the player doesn’t remember if a particular card in his hand was a club or a spade, or a diamond or a heart, respectively. (b) exp. 1 sample trial. answer options: yes/no. the epistemic state was as illustrated across all trials. (c) exp. 1 trial summary. all participants saw all trials, in random order. env pol modtype (comp, sup) x modmon (ue, de) = mod decl pos i have mod 3 [suit]. neg i don’t have mod 3 [suit] antcond pos if you have mod 3 [suit], then we have something in common. neg if you don’t have mod 3 [suit], then we have something in common. restuniv pos everyone who has mod 3 [suit] has something in common with me. neg everyone who doesn’t have mod 3 [suit] has something in common with me. (d) exp. 1 raw means and their associated 95% binomial cis. n = 96. figure 1: exp. 1 (a) instructions, (b) sample trial, (c) trial summary, and (d) raw results. proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 215 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) exp. 1 predicted effects – general (abridged, but listing all the significant contrasts). estimate std. error z value pr(>|z|) (intercept) 6.30 0.93 6.795 < 0.0001 modmonde -3.58 0.90 -3.978 0.0001 modtypesup -0.59 1.11 -0.532 0.5949 polneg -3.77 0.92 -4.103 < 0.0001 envrestuniv -1.68 0.95 -1.759 0.0785 modmonde:polneg 2.67 0.97 2.746 0.0060 modtypesup:polneg -1.69 1.17 -1.442 0.1492 (b) exp. 1 predicted effects – between modifier types (separated by modifier monotonicity). env pol mod or ci z p decl pos morethan-atleast 1.80 [0.13, 25.63] 0.532 1.0000 decl pos lessthan-atmost 2.02 [0.61, 6.64] 1.413 0.1576 decl neg morethan-atleast 9.73 [3.29, 28.83] 5.016 < 0.0001 decl neg lessthan-atmost 40.68 [12.66, 130.78] 7.597 < 0.0001 antcond pos morethan-atleast 0.84 [0.04, 19.27] -0.135 1.0000 antcond pos lessthan-atmost 2.13 [0.76, 5.96] 1.768 0.1540 antcond neg morethan-atleast 2.86 [1.13, 7.19] 2.722 0.0065 antcond neg lessthan-atmost 4.56 [1.77, 11.72] 3.842 0.0001 restuniv pos morethan-atleast 1.42 [0.24, 8.37] 0.477 1.0000 restuniv pos lessthan-atmost 3.56 [1.32, 9.62] 3.054 0.0068 restuniv neg morethan-atleast 3.59 [1.45, 8.88] 3.385 0.0014 restuniv neg lessthan-atmost 5.32 [2.10, 13.53] 4.294 < 0.0001 (c) exp. 1 predicted effects – within modifier types (separated by modifier monotonicity). env pol mod or ci z p decl-antcond neg morethan 3.37 [1.09, 10.42] 2.577 0.0199 decl-restuniv neg morethan 4.19 [1.38, 12.73] 3.092 0.0060 decl-antcond neg lessthan 5.07 [1.94, 13.25] 4.045 0.0002 decl-restuniv neg lessthan 3.78 [1.47, 9.74] 3.371 0.0015 decl-antcond neg atleast 0.99 [0.43, 2.29] -0.030 0.9760 decl-restuniv neg atleast 1.55 [0.68, 3.52] 1.272 0.5521 decl-antcond neg atmost 0.57 [0.19, 1.67] -1.253 0.4202 decl-restuniv neg atmost 0.50 [0.17, 1.43] -1.587 0.3373 table 1: exp. 1 model results. y ∼modmon * modtype * pol * env + (1 + (modmon + modtype + pol + env) | participant). (statistically significant results in gray.) over v2-3. however, this contrast was much larger in decl-neg than in antcond/restunivneg, due to the fact that, while cmns degrade between these conditions, smns do not. this is inexplicable on v1, where smns in antcond/restuniv-neg are expected to be even worse, just like cmns, but consistent with v2-3, where smns are actually expected to improve. we also failed to find a general penalty for smns. altogether, this suggests that the specific interaction of smns with negation is explained not by v1 but rather by v2-3. however, we also found a general penalty for the modifier monotonicity being de and for negation, and for their interaction. this suggests that v1 always plays a general role too. the fact that we also found a general trend for a penalty for the environment being the restriction of a universal is surprising on all of v1-3. proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 216 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3. experiment 2. 3.1. goal. to check what happens to smns in antcond/restuniv-pos/neg if we vary the pragmatic polarity of the predicate in the continuation, as a way to test the basic insight behind v2 that the key to smns in antcond/restuniv lies with questions of evaluativity. if v2 is correct, we expect this manipulation to make a big difference. in particular, an smn with a pragmatically neutral predicate in antcond/restuniv-pos should be fine with a positive continuation, and in antcond/restuniv-neg should be fine with a negative continuation. 3.2. methods. task: similar to exp. 1, adapted to support minimal manipulations of pragmatic predicate polarity in the continuation. see figure 2a. sample trial: see figure 2b. trial summary: each participant saw 32 trials, obtained by crossing the following factors: env = embedding environment (antcond = antecedent of conditional; restuniv = restriction of universal); pol1 = polarity of env (pos = positive, neg = negative); pol2 = pragmatic polarity of predicate in the continuation (pos = positive; neg = negative); and modtype x modmon = mod (as before). see figure 2c. participants: 45, of which 5 excluded (same answer in all trials). 3.3. results. for the raw means see figure 2d and for model results see tables 2a-2c. general effects: for both environment types, and for both cmns and smns, ratings were significantly lower for pol1=neg, the interaction of modmon=de, modtype=sup, and pol2=neg, the interaction of all the previous factors with env=restuniv, and the interaction of all the previous factors with pol1=neg. see table 2a. contrasts between modifier types: compared to its cmn counterpart more than, the ue smn at least: in pos-pos was similar; in pos-neg was similar in antcond but worse in restuniv; in neg-pos was worse; and in neg-neg was similar. compared to its cmn counterpart less than, the de smn at most was worse in every condition, although across the two environment types the magnitude of the contrast varied as follows: pos-neg > pos-pos > neg-pos > neg-neg. see table 2b. contrasts within modifier types: ratings were high for all the modifiers in pos-pos (with the notable exception of at most in restuniv), and generally remained so in pos-neg also—except for at most, which degraded dramatically (including in the case of restuniv). and ratings were lower for all the modifiers in neg-pos, and generally remained so in neg-neg also—except for at least, which improved dramatically. see table 2c. 3.4. discussion. ratings for smns in antcond/restuniv-pos/neg were affected by the pragmatic polarity of the continuation, both at least and at most being sometimes a lot better or a lot worse as a result of it. this is completely surprising on v1 or v3 but precisely as expected on v2. however, the patterns for at least and at most diverged vastly, and neither followed exactly the expectations from v2. this suggests that smns in antcond/restuniv-pos/neg are indeed sensitive to notions of evaluativity, as expected on v2, but the monotonicity of the modifier also makes a (big) difference. we also found a general penalty for negation, which again supports a general role of v1, and also a small penalty for the environment being the restriction of a universal (possibly specific to smn though the trend is not clear), which is again surprising on all of v1-3. we did not find a general penalty for the modifier being de, though this is likely still there, being simply obscured by the strong interaction of smn modifier monotonicity with the two kinds of polarity. proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 217 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) exp. 2 task. in this survey you will answer questions about a group of friends playing a game. at the beginning of the game each player gets dealt a hand of seven cards. they are not allowed to see their own cards but they are allowed to take a quick look at their neighbor’s hand. they try to remember their neighbor’s hand as well as they can because in the next step they have to come up with a rule that would make that neighbor (and possibly other players too) lose or win. you will see what a player remembers about their neighbor’s hand and the rule they make up, then you will be asked if you think the other players will understand what they said. note, we’re not asking you if it is a good rule or a bad rule, but whether it is a rule that is going to be understandable for the other players to follow. note: a or a means that the player doesn’t remember if a particular card in his hand was a club or a spade, or a diamond or a heart, respectively. (b) exp. 2 sample trial. answer options: yes/no. the epistemic state was as illustrated across all trials. (c) exp. 2 trial summary. all participants saw all trials, in random order. env pol1 pol2 modtype (comp, sup) x modmon (ue, de) = mod antcond pos pos/neg if you have mod 3 [suit], you win/lose. neg pos/neg if you don’t have mod 3 [suit], you win/lose. restuniv pos pos/neg everyone who has mod 3 [suit] wins/loses. neg pos/neg everyone who doesn’t have mod 3 [suit] wins/loses. (d) exp. 2 raw means and their associated 95% binomial cis. n = 40. figure 2: exp. 2 (a) instructions, (b) sample trial, (c) trial summary, and (d) raw results. proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 218 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) exp. 2 predicted effects – general (abridged, but listing all the significant contrasts). est. std. error z value pr(>|z|) (intercept) 6.07 1.21 5.028 < 0.0001 modmonde -1.87 1.30 -1.436 0.1509 modtypesup -1.33 1.30 -1.023 0.3063 pol1neg -4.25 1.19 -3.585 0.0003 pol2neg -0.78 1.23 -0.630 0.5288 modmonde:modtypesup:pol2neg -3.77 2.10 -1.793 0.0730 modmonde:modtypesup:pol2neg:envrestuniv 6.19 3.02 2.050 0.0404 modmonde:modtypesup:pol1neg:pol2neg:envrestuniv -6.43 3.48 -1.849 0.0644 (b) exp. 2 predicted effects – between modifier types (separated by modifier monotonicity). env pol1 pol2 mod or ci z p antcond pos pos morethan-atleast 3.78 [0.20, 69.85] 1.023 0.5451 antcond pos pos lessthan-atmost 13.15 [1.45, 119.61] 2.616 0.0089 antcond pos neg morethan-atleast 1.76 [0.12, 26.56] 0.467 0.6402 antcond pos neg lessthan-atmost 264.82 [29.54, 2374.44] 5.701 < 0.0001 antcond neg pos morethan-atleast 4.80 [1.18, 19.54] 2.506 0.0244 antcond neg pos lessthan-atmost 10.71 [2.28, 50.38] 3.434 0.0012 antcond neg neg morethan-atleast 0.67 [0.12, 3.64] -0.532 0.5946 antcond neg neg lessthan-atmost 6.90 [1.48, 32.22] 2.809 0.0099 restuniv pos pos morethan-atleast 4.00 [0.24, 68.07] 1.097 0.5451 restuniv pos pos lessthan-atmost 96.56 [6.66, 1400.93] 3.830 0.0003 restuniv pos neg morethan-atleast 48.06 [2.34, 988.59] 2.870 0.0082 restuniv pos neg lessthan-atmost 102.38 [13.90, 753.84] 5.196 < 0.0001 restuniv neg pos morethan-atleast 2.91 [0.75, 11.26] 1.769 0.0769 restuniv neg pos lessthan-atmost 4.74 [1.09, 20.65] 2.370 0.0178 restuniv neg neg morethan-atleast 0.39 [0.09, 1.62] -1.480 0.2776 restuniv neg neg lessthan-atmost 3.75 [0.79, 17.86] 1.895 0.0581 (c) exp. 2 predicted effects – within modifier types (separated by modifier monotonicity). env pol1 pol2 mod or ci z p antcond pos pos-neg morethan 2.17 [0.14, 34.51] 0.630 1.0000 antcond pos pos-neg lessthan 1.65 [0.15, 17.65] 0.471 0.6374 antcond pos pos-neg atleast 1.01 [0.07, 14.96] 0.010 0.9920 antcond pos pos-neg atmost 33.16 [6.61, 166.30] 4.867 < 0.0001 antcond neg pos-neg morethan 0.82 [0.19, 3.46] -0.310 0.7568 antcond neg pos-neg lessthan 1.35 [0.41, 4.48] 0.557 0.5775 antcond neg pos-neg atleast 0.11 [0.03, 0.50] -3.280 0.0021 antcond neg pos-neg atmost 0.87 [0.17, 4.33] -0.197 0.8614 restuniv pos pos-neg morethan 0.42 [0.02, 11.43] -0.590 1.0000 restuniv pos pos-neg lessthan 3.53 [0.21, 58.42] 1.008 0.6271 restuniv pos pos-neg atleast 5.02 [0.54, 46.87] 1.620 0.2103 restuniv pos pos-neg atmost 3.74 [1.06, 13.22] 2.345 0.0190 restuniv neg pos-neg morethan 1.91 [0.52, 7.04] 1.114 0.5307 restuniv neg pos-neg lessthan 2.19 [0.62, 7.72] 1.399 0.3235 restuniv neg pos-neg atleast 0.26 [0.07, 0.93] -2.360 0.0183 restuniv neg pos-neg atmost 1.73 [0.36, 8.28] 0.788 0.8614 table 2: exp. 2 model results. y ∼modmon * modtype * pol1 * pol2 * env + (1+ (modmon + modtype + pol1 + pol2 + env)|participant). (statistically significant results in gray.) proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 219 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4. experiment 3. 4.1. goal. to check what happens to smns under [de-environment]-pos/neg if we vary the de environment from antcond to the scope of a negated factive, as a way to test the basic insight behind v3 that the key to smns in antcond/restuniv lies with questions of monotonicity. if v3 is correct, we expect this manipulation to make no difference. in particular, an smn under pos/neg should be the same whether further embedded in antcond or a negated factive, as both antcond and a negated factive define a de environment with an ue presupposition. 4.2. methods. task: similar to exp. 1, adapted to support minimal manipulations of the matrix-level de environment. see figure 3a. sample trial: see figure 3b. trial summary: 16 trials, obtained by crossing the following factors: env = matrix embedding environment (antcond = antecedent of conditional that here also contains a factive, matrixneg = scope of matrix negation which here also contains a factive); pol = polarity of the embedded clause (pos = positive, neg = negative); and modtype x modmon = mod (as before). see figure 3c. participants: 45, none excluded. 4.3. results. for the raw means see figure 3d and for model results see tables 3a-3c. general effects: for all environment types, and for both cmns and smns, ratings were significantly lower for modmon=de and pol=neg. see table 3a. contrasts between modifier types: for all environment types, for the same level of monotonicity, smns (a) in pol=pos are similar to cmns but in pol=neg are sometimes worse, this effect being clearly detected for at least, especially in matrixneg. see table 3b. contrasts within modifier types: ratings for all the modifiers degrade somewhat between antcond and matrixneg, but for smns this effect was sometimes found significant, being detected in pol=pos for at most and in pol = neg for both at least and at most. see table 3c. 4.4. discussion. ratings for smns in pos/neg had similar patterns regardless of the choice of a matrix de environment. this is completely surprising on v1 or v2 but precisely as expected on v3. the biggest point of difference comes from antcond-pos and matrixneg-pos. on v1 and v2, on which these conditions are merely more complex variants of antcond-pos and decl-neg in exp. 1, smns were expected to be fine in antcond-pos but very degraded in matrix-neg, contrary to what was observed. however, on v3, on which the presence of the factive does not just make these conditions more complex than their exp. 1 counterparts but in fact fundamentally distinguishes matrixneg-pos from decl-neg, endowing it with an ue presupposition, smns are expected to be fine in both antcond-pos and matrixneg-pos, just as observed. this validates the insight from v3 that smns in antcond(/restuniv) care about monotonicity. however, although similar across the two matrix de environments, smns still degraded slightly from antcond to matrixneg. this suggests that, while smns in antcond(/restuniv)-pos/neg are indeed sensitive to a deeper notion of monotonicity, as expected on v3, the specific identity of the de operator (and possibly also whether it is actual negation) also makes a (small) difference. (note: this brings to mind the finding from exps. 1-2 that, although similar to antcond, restuniv was always slightly worse.) in addition to this, we also found a general penalty for the modifier monotonicity being de and for the embedded polarity being negative, suggesting yet again that v1 always plays a general role also. proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 220 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) exp. 3 task. in this survey you will consider a commentator for a televised card-playing game, and answer questions about how understandable the commentator is. at the beginning of the game each player gets dealt seven cards, two of which are hidden. then in each round some rule is issued, and players can choose whether or not to bet on their own hand. a commentator, who knows what the hidden cards are for each player, discusses the player’s move. you will see a player’s hand and the commentator’s comment, then you will be asked if you think the viewers will understand what the commentator said. note: in the hands that you will see, cards with a white background such as represent cards that are visible to the player, while cards with a grey background such as represent hidden cards, that is, cards that are not visible to the player but visible to the commentator. (b) exp. 3 sample trial. answer options: yes/no. the epistemic state was as illustrated across all trials. (c) exp. 3 trial summary. all participants saw all trials, in random order. env pol modtype (comp, sup) x modmon (ue, de) = mod matrixneg pos [name] doesn’t know that s/he has mod 3 [suit] neg [name] doesn’t know that s/he doesn’t have mod 3 [suit] antcond pos if [name] knew that s/he has mod 3 [suit], s/he would bet differently neg if [name] knew that s/he doesn’t have mod 3 [suit], s/he would bet differently (d) exp. 3 raw means and their associated 95% binomial cis. n = 45. figure 3: exp. 3 (a) instructions, (b) sample trial, (c) trial summary, and (d) raw results. proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 221 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) exp. 3 predicted effects – general (abridged, but listing all the significant contrasts). estimate std. error z value pr(>|z|) (intercept) 4.30 0.96 4.464 < 0.0001 modmonde -2.34 0.94 -2.480 0.0131 modtypesup -0.26 1.12 -0.233 0.8161 polneg -2.15 0.99 -2.186 0.0288 envmatrixneg -0.52 0.96 -0.539 0.5898 modmonde:modtypesup -0.03 1.22 -0.021 0.9834 modmonde:polneg -0.91 1.10 -0.828 0.4078 modtypesup:polneg -0.83 1.22 -0.677 0.4987 modmonde:envmatrixneg 0.07 1.09 0.067 0.9468 modtypesup:envmatrixneg -0.64 1.30 -0.497 0.6190 polneg:envmatrixneg -0.32 1.11 -0.290 0.7716 modmonde:modtypesup:polneg 1.58 1.47 1.077 0.2816 modmonde:modtypesup:envmatrixneg -0.21 1.51 -0.139 0.8891 modmonde:polneg:envmatrixneg 0.70 1.38 0.508 0.6112 modtypesup:polneg:envmatrixneg 0.25 1.51 0.169 0.8661 modmonde:modtypesup:polneg:envmatrixneg -0.61 1.89 -0.321 0.7479 (b) exp. 3 predicted effects – between modifier types (separated by modifier monotonicity). env pol mod or ci z p antcond pos morethan-atleast 1.30 [0.10, 16.06] 0.233 0.8161 antcond pos lessthan-atmost 1.33 [0.30, 5.96] 0.428 0.6687 antcond neg morethan-atleast 2.96 [0.69, 12.71] 1.672 0.0946 antcond neg lessthan-atmost 0.62 [0.16, 2.37] -0.793 0.5065 matrixneg pos morethan-atleast 2.47 [0.34, 17.98] 1.023 0.6131 matrixneg pos lessthan-atmost 3.13 [0.84, 11.63] 1.946 0.1033 matrixneg neg morethan-atleast 4.38 [1.17, 16.40] 2.505 0.0245 matrixneg neg lessthan-atmost 2.09 [0.49, 8.85] 1.143 0.5065 (c) exp. 3 predicted effects – within modifier types (separated by modifier monotonicity). env pol mod or ci z p antcond-matrixneg pos morethan 1.68 [0.20, 14.37] 0.539 0.5898 antcond-matrixneg pos lessthan 1.56 [0.38, 6.38] 0.706 0.9598 antcond-matrixneg pos atleast 3.19 [0.35, 29.04] 1.178 0.2386 antcond-matrixneg pos atmost 3.66 [1.02, 13.11] 2.282 0.0450 antcond-matrixneg neg morethan 2.32 [0.56, 9.64] 1.320 0.3739 antcond-matrixneg neg lessthan 1.07 [0.29, 3.89] 0.111 0.9598 antcond-matrixneg neg atleast 3.42 [1.10, 10.66] 2.426 0.0305 antcond-matrixneg neg atmost 3.57 [0.90, 14.23] 2.065 0.0450 table 3: exp. 3 model results. y ∼modmon * modtype * pol * env + (1 +(modmon + modtype + pol + env) |participant). (statistically significant results in gray.) proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 222 https://doi.org/10.3765/elm https://www.elm-conference.net/ 5. conclusion. in contrast to cmns, smns are reported to degrade under negation, though not in all de environments. the literature offers four views: this contrast (v0) does not exist, or it exists and it is about (v1) processing cost, (v2) monotonicity and evaluativity (cohen & krifka 2014), or (v3) monotonicity alone (spector 2015, mihoc 2020). in this paper we found the contrast does exist, pace v0, but is richer than processing cost, pace v1, the interaction between smns under negation being multiply negotiable, as on v2-3. for the reasons discussed above and others,1 we believe the basic interaction of smns with de environments is best captured by v3 (spector 2015, mihoc 2020). at the same time, we conclude that v3 must be enriched to capture the evaluativity effect that inspired v2 (for an attempt see mihoc to appear) and the general fixed costs of various de operators that inspired v1. references alexandropoulou, stavroula. 2018. on the pragmatics of numeral modifiers: the availability and time course of variation, ignorance and indifference inferences. utrecht: lot. clark, herbert h. & william g. chase. 1972. on the process of comparing sentences against pictures. cognitive psychology 3(3). 472–517. 10.1016/0010-0285(72)90019-9. cohen, ariel & manfred krifka. 2014. superlative quantifiers and meta-speech acts. linguistics and philosophy 37(1). 41–90. 10.1007/s10988-014-9144-x. cremers, alexandre & emmanuel chemla. 2017. experiments on the acceptability and possible readings of questions embedded under emotive-factives. natural language semantics 25(3). 223–261. 10.1007/s11050-017-9135-x. geurts, bart & rick nouwen. 2007. at least et al.: the semantics of scalar modifiers. language 533–559. 10.1353/lan.2007.0115. mayr, clemens. 2013. implicatures of modified numerals. in ivano caponigro & carlo cecchetto (eds.), from grammar to meaning: the spontaneous logicality of language, 139–171. cambridge: cup. 10.1017/cbo9781139519328. mihoc, teodora. 2020. ignorance and anti-negativity in the grammar: or/some and modified numerals. in the annual meeting of the north east linguistic society (nels) 50, 197–210. mihoc, teodora. to appear. modified numerals and polarity sensitivity: between o(nly)da and e(ven)sa. in sinn und bedeutung (sub) 25, tba. nouwen, rick, stavroula alexandropoulou & yaron mcnabb. 2019. experimental work on the semantics and pragmatics of modified numerals. in chris cummins & napoleon katsos (eds.), handbook of experimental semantics and pragmatics, oxford: oup. 10.1093/oxfordhb/9780198791768.013.15. spector, benjamin. 2015. why are class b modifiers global ppis? handout for talk at workshop on negation and polarity, february 8-10, 2015, the hebrew university of jerusalem. wason, peter c. 1961. response to affirmative and negative binary statements. british journal of psychology 52(2). 133–142. 10.1111/j.2044-8295.1961.tb00775.x. 1further reasons include: smns are degraded not just under negation but under many other de operators (e.g., without), suggesting v1 is too limited; smns are bad under negation even with a stereotypically positive property (e.g., jo didn’t solve # at least 3 problems), suggesting the v2 solution for negation is incomplete, and all the de environments where smns are fine seem to contain an ue presupposition (e.g., the scope of only), suggesting the v2 solution for conditionals / universals is missing a generalization. proceedings of elm 1: 212-223, 2021 teodora mihoc and kathryn davidson: superlative-modified numerals and negation: a multiply negotiable cost. 223 https://doi.org/10.3765/elm https://www.elm-conference.net/ discourse behavior of possessives reflects the importance of interpersonal relationships jesse storbeck & elsi kaiser* abstract. nominal possessive constructions (e.g. sam’s car) present a challenge for theories of discourse since, unlike simpler nominal phrases (e.g. a car), they explicitly refer to two entities, not just one. research on the discourse prominence of these two referents has been limited in scope and produced contradictory findings. we use a sentence continuation experiment to investigate the prominence of possessions as a function of their animacy. we find that possessed animates (e.g. her butler) are especially prominent. their privileged status in discourse may relate to non-linguistic theories on the importance of interpersonal relationships. keywords. psycholinguistics; discourse; possessives; sentence continuation 1. introduction. a widely held view about discourse-level representation and processing is that referents in a given discourse vary in prominence (alternatively, ‘salience’) and that the prominence of referents changes over time as the discourse unfolds (e.g. ariel, 1988; van den broek et al., 1996). many factors contribute to a referent’s prominence, but two of the most influential and dependable predictors are grammatical role and animacy; specifically, subjects tend to be more prominent than objects (e.g. chafe, 1976; crawley et al., 1990), and animates tend to be more prominent than inanimates (e.g. bock et al., 1992; dahl & fraurud, 1996). however, to the best of our knowledge, prior work has tended to overlook a frequent grammatical structure with potential to inform current theories of discourse: nominal possessive constructions (e.g. sam’s car, sam’s doctor). unlike simpler nominals (a car, the doctor), nominal possessives reference two entities, a possessor (sam) and a possession (car/doctor) within the same noun phrase. since most previous work investigating the discourse prominence of referents has focused on nominal phrases containing a single referent, the lack of research on possessives brings up two main questions. firstly, how are the two entities represented in discourse? do they simply behave like independent discourse referents, or are their representations linked somehow? the latter seems to be necessary if we hope to encode the possessive relationship in the discourse. secondly, if the two discourse referents are linked, does the link differ across different types of semantic possession relations? this question is particularly notable for languages like english, which use the same syntactic configuration to express a variety of relations between the possessor and possession, such as ownership (e.g. sam’s book), part-whole (e.g. sam’s arm), and kinship (e.g. sam’s mother). in contrast, other languages, such as maltese, use different morphosyntactic mechanisms to express different types of possession relations. for example, maltese id ‘hand’ and ktieb ‘book’ require different morphosyntax to form possessives: id-i ‘my hand’ (part-whole relation) and il-ktieb tiegħ-i ‘my book’ (ownership relation) (haspelmath, 2017). before proceeding further, we should note that the type of nominal possessive exemplified by sam’s car is sometimes known as an ‘s-genitive’; however, english also realizes nominal possessives with the ‘of-genitive’ structure (e.g. the floor of the bedroom, the capital of califor * the authors gratefully acknowledge the audience at the experiments in linguistic meaning conference, virtually hosted by the university of pennsylvania, for their useful comments and feedback. we also thank the members of the usc language processing lab for their input. additional thanks are due to madeline rouse for assisting with the data annotation. authors: jesse storbeck, university of southern california (jstorbec@usc.edu) & elsi kaiser, university of southern california (emkaiser@usc.edu). proceedings of elm 1: 261-272, 2021 c©2021 jesse storbeck and elsi kaiser published by the lsa with permission of the author(s) under a cc by license. 261 https://doi.org/10.3765/elm https://www.elm-conference.net/ nia). this study focuses on english s-genitives for multiple reasons. firstly, animate possessors exhibit a broad range of semantic possession types—key for investigating how different possession relations affect discourse-level representation—yet animates are typically dispreferred as possessors in of-genitives (e.g. rosenbach 2002). additionally, our work will involve pronouns, which can be ungrammatical or degraded as possessors in of-genitive constructions (e.g. the palace of the king vs. *the palace of him; the roof of the house vs. ?the roof of it) (e.g. rosenbach 2002). the present study investigates how nominal possessive constructions are encoded in discourse and how different semantic possession relations affect the discourse-level representations for the referents involved in possessives. we begin by reviewing prior work on the discourse properties of possessives, which comes from applications of centering theory (e.g. grosz et al., 1995) (section 1.1). we also summarize some relevant research on interpersonal relationships, which possessives explicitly denote when both the possessor and possession are human (section 1.2). based on existing work, we then propose three hypotheses about the discourse prominence of possessed referents (section 1.3). in section 2, we test our hypotheses using a sentence continuation task. in section 3, we present our results, which suggest that possessed animate referents are especially prominent in discourse. we conclude (section 4) in favor of theory that links the general cognitive importance of interpersonal relationships with the apparently privileged discourse representations of their linguistic realizations. 1.1. previous approaches based on centering theory. while the discourse behavior of nominal possessives has generally received little attention in psycholinguistic literature, some researchers have approached this issue within the framework of centering theory (grosz et al., 1995). centering theory is a model of discourse coherence which seeks to capture patterns in transitions from one utterance to another (e.g. what makes a particular sequence of utterances more or less coherent as a discourse). centering theory is particularly relevant to the present study because a key component of its algorithm is ranking referents according to their prominence. given the frequency of nominal possessive constructions and the fact that most work within centering theory has been corpus-based, it is not surprising that this framework in particular has had to address nominal possessive constructions. within centering theory, researchers have argued for two competing accounts of the relative prominence of possessors and possessions. on the one hand, chae (2003) follows the complex np assumption of walker & prince (1996), which posits that the discourse referents realized within a complex noun phrase—of which possessives are one type—are ranked from most to least prominent based on their left-to-right order. therefore, possessors in s-genitives are more prominent than their possessions (e.g. in sam’s car, sam would outrank car in prominence). on the other hand, di eugenio (1998) proposes (for possessives involving animate possessors) that the ranking depends on the animacy of the possession; animate possessions immediately outrank their possessors, while inanimate possessions are ranked immediately below their possessors. accordingly, in sam’s car, sam would still outrank car, but in sam’s doctor, doctor would outrank sam. this theory dovetails with the well-documented cross-linguistic finding in the broader discourse literature that animate referents tend to be more prominent than inanimate ones (e.g. bock et al., 1992; dahl & fraurud, 1996; dahl, 2008). for instance, animate referents tend to appear before inanimates and in subject position (e.g. branigan et al., 2007; prat-sala & branigan, 2000). additionally, animacy influences the linguistic form with which a referent is mentioned, with animates being overall more likely to be pronominalized than inanimates (arnold & griffin, 2007; fukumura & van gompel, 2011). proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 262 https://doi.org/10.3765/elm https://www.elm-conference.net/ notwithstanding the debate within centering theory on how to deal with nominal possessives, both of these competing accounts leave an important question unanswered: how does being possessed affect a referent? for instance, if we compare possessive constructions like sam’s car to simpler noun phrases like the car or a car, will there be any difference in the discourse prominence of car? given that a key goal of theories of discourse representation is to shed light on the factors that modulate a referent’s prominence, an understanding of the effect of possession itself is needed in a theory of the discourse behavior of nominal possessives. 1.2. the importance of interpersonal relationships. when the possessor and possession in a possessive construction are both animate and human, the construction explicitly denotes an interpersonal relationship. such relationships may be permanent (e.g. sam’s mother, the teacher’s son) or more transitory (e.g. their neighbor, the accountant’s boss) and may vary with respect to the formality of the relationship’s definition (e.g. strictly biological relations vs. gradient social relations). nevertheless, in all possessives where both possessor and possession are human, the two individuals stand in some kind of socially salient and not-entirely transient relation to each other (i.e. an interpersonal relationship). in non-linguistic research, interpersonal relationships have been shown to be critical for human health and well-being (e.g. cacioppo & hawkley, 2009; eisenberger & cole, 2012). for instance, social connectedness has been linked to longevity and disease resistance (e.g. miller et al., 2009; holt-lunstad et al., 2010), while loneliness is correlated with depression and cognitive decline (e.g. tilvis et al., 2004; cacioppo et al., 2006). in short, this literature makes clear that people with stronger connections to others tend to live longer, happier lives, while socially isolated individuals tend to suffer a variety of negative consequences. furthermore, prior research on memory shows that people tend to form stronger memory representations for animate entities than for inanimates (e.g. nairne et al., 2013; bonin et al., 2014). this observation has led some researchers to theorize that that better memory for animates arose from the evolutionary importance of identifying threats, mates, and social groups, most of which entail interpersonal relationships (e.g. nairne, 2010; vanarsdall et al. 2013). based on such findings, the human mind appears to be especially attuned to interpersonal relationships. therefore, one might expect nominal possessives that express interpersonal relationships to also be privileged in our mental representations, relative to other kinds of possessive relations. the experiment reported in this paper explores whether these general cognitive tendencies have any effect on a referent’s discourse prominence, as reflected by language production. before turning to the experiment itself, we next outline three possible hypotheses about the discourse representations of possessed referents. 1.3. three hypotheses concerning the discourse representations of possessions. based on prior work, we propose and test three novel hypotheses about the discourse representations of possessed referents. importantly, the first two hypotheses are not mutually exclusive; the third hypothesis supersedes the previous two, while not necessarily invalidating their underlying logic. the first hypothesis is the animacy hypothesis. animate referents are widely viewed as more prominent in discourse and memory, since they tend to be mentioned earlier (e.g. bock et al., 1992; dahl, 2008), are more frequently pronominalized (e.g. arnold & griffin, 2007; fukumura & van gompel, 2011), and persist in memory (in both a linguistic and domain-general sense) more than inanimates (e.g. nairne et al., 2013; bonin et al., 2014). therefore, the animacy hypothesis proposes that animate possessions (e.g. sam’s doctor) are more prominent in discourse than inanimate ones (e.g. sam’s car) for the same reasons and to the same extent that simpler nominals exhibit animacy effects (the/a doctor vs. the/a car). proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 263 https://doi.org/10.3765/elm https://www.elm-conference.net/ the second hypothesis is the possessive hypothesis. nominal possessives like sam’s car are referentially and semantically more complex than simpler nominals due the presence of an additional referent and the link between possessor and possession. since increased representational complexity has been shown to promote retrieval from memory (e.g. fisher & craik, 1980; hofmeister, 2011; karimi & ferreira, 2016; troyer et al., 2016), we might expect possessions to be more prominent than simpler nominals on the discourse level. therefore, according to the possessive hypothesis, car in sam’s car would be more prominent than car in the/a car. as stated previously, the animacy hypothesis and the possessive hypothesis are not mutually exclusive, and were they both to hold, the result would be that possessed animates are more prominent than either possessed inanimates or non-possessed animates; however, there might be another reason—irrespective of the combination of the previous two hypotheses—to expect increased prominence for possessed animates: their explicit denotation of interpersonal relationships (e.g. sam’s mother). as discussed in section 1.2, it seems reasonable to assume, based on prior work, that humans’ mental representations of interpersonal relationships are cognitively privileged (e.g. cacioppo & hawkley, 2009; eisenberger & cole, 2012). based on findings within the social psychology, health, and memory literatures, we might additionally suspect that their linguistic realizations are privileged on the discourse level. therefore, our third hypothesis is the interaction hypothesis, whereby possessed animates are especially prominent in discourse—in excess of any additive effects of animacy and possession—due to the domain-general cognitive significance of interpersonal relationships. accordingly, we would expect doctor in sam’s doctor to be especially prominent, in excess of any additive effects of animacy and possession status. 2. testing our hypotheses with a sentence continuation experiment. our examination of existing work has illustrated that there is still a significant gap in theories of discourse representation and processing concerning the nature of possessives. we therefore seek to investigate how possession affects a referent and whether different semantic possession relations modulate that effect. to this end, we have proposed three novel hypotheses concerning the representation of possessed referents, which we hope to support or reject using a sentence continuation experiment. the sentence continuation task is commonly viewed as providing a measure of referents’ prominence. it builds on the common assumption that a referent’s prominence is positively correlated with its likelihood of subsequent mention in the discourse (e.g. givón, 1983; arnold, 2001; kehler et al., 2008; kaiser, 2009; kehler & rohde, 2013). furthermore, existing discourse theory posits that the most prominent referents tend to appear as grammatical subjects (e.g. chafe, 1976; gordon et al., 1993; stevenson et al., 1994). thus we can measure which referents are most prominent by analyzing how often they are selected as continuation subjects (e.g. givón, 1983; ariel, 1988; stevenson et al., 1994; arnold, 2001). we also analyze mentions of entities throughout the continuation sentence as a further measure of discourse prominence, as well as looking at the linguistic form of mentions as a measure of entities’ conceptual accessibility (e.g. kaiser, 2009). 2.1. participants. data from 40 native english speakers is presented for analysis. all participants were 18 years of age or older and were recruited from amazon mechanical turk. all of the participants included in the analysis reported being born in the united states and identified as native speakers of english. all reported normal or corrected-to-normal vision and normal hearing. proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 264 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.2. design and materials. the experiment had 56 items: 24 targets and 32 fillers. all target items followed the frame: [name] [nonce verb in simple past tense] [a vs. his/her] [animate vs. inanimate] (e.g. jessica rulked an electrician; see example in table 1). indefinite animate possessed animate indefinite inanimate possessed inanimate jessica rulked an electrician. jessica rulked her electrician. jessica rulked a chandelier. jessica rulked her chandelier. table 1: an example target item (1 of 24) in each of the four conditions we manipulated (i) possession: whether the object was possessed or indefinite and (ii) animacy: whether the direct object (which was the possession in the possessed conditions) was animate (and human) or inanimate. animate objects were all human role nouns (e.g. chauffeur, florist, housekeeper, stockbroker); inanimate objects were all concrete alienable possessions (e.g. jacket, stereo, toaster, umbrella). to minimize the potential for referential ambiguity in participants’ continuations, the names (i.e. the prompt subjects) were all unambiguous with respect to their typically associated gender and matched for each item with human role noun objects which were stereotypically biased toward the opposite gender. the order of the genders in target items with animate objects was counterbalanced, as was the gender of the subjects overall. for each target item, animate objects were paired with inanimates that matched as closely as possible in lexical frequency, syllable length, and character length. we chose to contrast possessive noun phrases with indefinites, rather than definites, because we found indefinites to sound more natural in the abbreviated contexts of our items; however, we intend to test definites as well in future work, since the givenness contexts which license definites are perhaps a more apt fit for possessives (e.g. gundel et al., 1993; barker, 2000). nonce verbs (e.g. blorned, chabbed, dasped, tammed) were used because they allowed the verb to remain constant within items, since predicates which naturally take animate objects are often unnatural with inanimate objects (and vice versa). the nonce verbs also limited the effects of verbal semantics on the continuations, perhaps from implicit causality (e.g. hartshorne & snedeker, 2013) or distributional biases toward animate or inanimate objects. the 2 × 2 (indefinite/possessive × animate/inanimate) design resulted in four conditions. conditions were distributed across four experimental lists in a latin square (i.e. within-subject and within-item design). thus, each participant saw six items per condition. targets and fillers on lists were interleaved in a pseudorandomized fashion. four additional lists were created by reversing the trial order of the original lists, for a total of eight experimental lists. 2.3. procedure. before beginning the study, participants were told that they would read prompt sentences and write one-sentence continuations. participants were instructed to make their continuations natural-sounding, not to copy-paste material from previous continuations or the prompts, and to limit their responses to a single complete sentence for each item. they completed three practice items (crucially without any possessive structures) and saw samples of acceptable continuations and unacceptable fragments. other than the instructions against fragments and copy-paste, participants were assured multiple times that there were no right or wrong answers. experimental items were presented on separate pages (see the example in figure 1). proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 265 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: an example target item as it appeared to participants 2.4. annotation of the continuation data. in order to analyze participants’ continuations, they were annotated by hand with respect to mentions of the preceding subject and object. additionally, we annotated the head noun of the subject noun phrase for every clause in the continuation. analyses below based on subject position are limited to the first independentclause subject in each continuation. accordingly, if the prompt was jennifer pranned her surgeon, and the participant wrote she asked him how long her surgery would last, the continuation would be coded as mentioning both the preceding subject and preceding object and mentioning the preceding subject as subject of the continuation. in the analyses below, a word was only counted as a mention of an entity if the head of the noun phrase referred to that entity; for example, for the prompt daniel zatted his jacket, an instance of his jacket in the continuation would only count as a mention of the preceding object (jacket) and not the preceding subject (daniel). 2.5. predictions. crucially, our three hypotheses introduced in section 1.3 make different predictions of the results. here we frame our predictions for the animacy hypothesis, possessive hypothesis, and interaction hypothesis in relation to mentions of the preceding object, since this is the locus of the experimental manipulation. notably, these predictions are broadly relevant for both the subject-position and entire-continuation analyses. we consider subject position to reflect a “winner take all” measure of discourse prominence, under the widely held assumption that the most prominent referents in an utterance tend to be realized as grammatical subjects. on the other hand, we expect that mentions across the entire continuation will be a correlated but more inclusive measure that may pick up on finer differences in prominence. as discussed in section 1.3, the animacy hypothesis posits that, due to animates’ general advantage in prominence over inanimates, animate possessions are therefore more prominent than inanimate ones. accordingly, the animacy hypothesis predicts that animate possessions (e.g. his nurse) will be more likely to be mentioned in continuations than inanimate possessions (e.g. his jacket), and that this difference will be comparable in magnitude to the difference we observe for the indefinites (e.g. a nurse vs. a jacket). in statistical terms, we therefore expect a main effect of animacy with respect to mentions of the preceding object (without making any commitments as to the effect of possession). the possessive hypothesis states that possessed referents are more prominent than simple noun phrases due to their referential and semantic complexity. accordingly, it predicts that the possessed preceding objects will be more likely to be mentioned than the indefinites. we would expect such an effect of possession to apply approximately equally to animates and inanimates (i.e. a main effect of possession in our statistical models). finally, the interaction hypothesis theorizes that animate possessions are exceptionally prominent in discourse due to their denotation of interpersonal relationships. therefore, we proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 266 https://doi.org/10.3765/elm https://www.elm-conference.net/ would expect possessed animates to receive an extra boost in likelihood of mention in participants’ continuations (i.e. a statistical interaction of animacy and possession). notably, animacy and possession could still contribute independently to the prominence of possessed animates following the logic of the animacy hypothesis and possessive hypothesis; however, the interaction hypothesis rests on observing a superadditive effect of animacy and possession together. 3. results. all statistical analyses used generalized linear mixed-effects models implemented with the r package lme4 (bates et al., 2015). all the models were fit to the binomial outcome of mentioning some entity in a certain position (1) or not (0). the independent variables animacy and possession were deviation coded (animate = 0.5, inanimate = -0.5; possessive = 0.5, indefinite = -0.5). we used the maximal random effects structure in each model that did not result in nonconvergence (barr et al., 2013). when a model’s random effects structure needed to be reduced due to nonconvergence, we prioritized the inclusion of random slopes for participants. 3.1. mentions in subject position. we first analyzed the subject position of continuations to see whether participants mentioned the preceding subject, object, or a third party not mentioned in the prompt. when a continuation contained multiple clauses, we analyzed the subject of the first independent clause. a visualization of these data is given below in figure 2. figure 2: does the subject of the continuation sentence refer back to the preceding subject, preceding object, or something else? (proportions are separated by condition) we fit two models which predicted the probability of the continuation subject referring (with any linguistic form) to the preceding subject (model 1) or preceding object (model 2). model 1 revealed an animacy effect, whereby the presence of an animate object in the prompt sentence significantly reduced the likelihood of the continuation subject referring to the prompt subject (i.e. the proper name) (p < 0.01). this main effect of animacy seems to reflect increased competition for prominence between the preceding subject and object when the object is animate. neither the main effect of possession nor its interaction with animacy were significant (p = 0.35 and p = 0.74, respectively). like the model for the preceding subject, model 2 (predicting the probability of mentioning the preceding object in subject position) also showed a significant main effect of animacy (p < 0.001). this finding illustrates that animate objects had a significantly greater chance of being 0.0 0.1 0.2 0.3 0.4 0.5 0.6 indefinite animate possessed animate indefinite inanimate possessed inanimate pr op or tio n of c on tin ua tio ns continuation subject refers to preceding... subject object other proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 267 https://doi.org/10.3765/elm https://www.elm-conference.net/ mentioned in subject position compared to inanimate objects, reflecting the well-known advantage in prominence that animacy confers. there was no significant main effect of possession (p = 0.30), but the model did show a significant interaction of possession and animacy (p = 0.04), which reflects a boost in the likelihood of mention for animate possessions. to further examine this interaction, we fit an additional model to test for a simple effect of possession within the animate conditions; this analysis revealed a significant simple effect of possession (p = 0.02), supporting the interpretation that preceding animate objects were more likely to be mentioned as continuation subjects when they were possessed. there was no corresponding simple effect of possession in the inanimate conditions (p = 0.43). this finding suggests that animate possessions are especially prominent in discourse and consequently supports the interaction hypothesis. 3.2. mentions in all positions. we turn now to the analysis of mentions in all positions of the continuations. a visualization of references anywhere in the continuation (with any linguistic form) is presented as figure 3. the statistical analysis proceeded in much the same way as the subject-position analysis, except the binomial outcome was mention of a particular entity across the entire continuation. again, we fit two models, one for the preceding subject (model 3) and one for the preceding object (model 4). figure 3: does the continuation sentence contain a reference to the preceding subject or object? (proportions are separated by condition) model 3 (predicting the likelihood of mentioning the preceding subject anywhere in the continuation) revealed no significant effects (all p > 0.50). this result contrasts with the corresponding subject-position analysis (model 1, section 3.1), which showed an animacy effect for mentions of the preceding subject. while the preceding subject was less likely in animate conditions to be mentioned as the subject of the continuation (as shown in figure 2), by the end of the continuation, all the conditions showed a similar likelihood of having mentioned the preceding subject somewhere (as shown in figure 3). model 4 (predicting the likelihood of mentioning the preceding object anywhere in the continuation) showed a significant effect of possession (p < 0.01) and, like the corresponding subject-position analysis (model 2, section 3.1), a significant interaction of possession and animacy (p < 0.01). this interaction shows that possessed animates were especially likely to be 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 indefinite animate possessed animate indefinite inanimate possessed inanimate pr op or tio n of c on tin ua tio ns continuation contains mention of preceding... subject object proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 268 https://doi.org/10.3765/elm https://www.elm-conference.net/ mentioned across the entire continuation. finally, we again sought to confirm this interaction by testing for a simple effect of possession within just the animate conditions; like the corresponding subject-position analysis in section 3.1, this follow-up model showed that possessed animates were significantly more likely than their indefinite counterparts to be mentioned in continuations (p < 0.001); however, no such simple effect was present for the inanimate conditions (p = 0.83). therefore, our results again support the interaction hypothesis and the theory that animate possessions enjoy special discourse prominence. 4. discussion. not all referents are equally prominent in comprehenders’ mental representations of discourse. we report an experiment investigating the discourse behavior of nominal possessives (e.g. sam’s car), which differ from simpler definite or indefinite nouns (e.g. a/the car) by referencing two entities: a possessor (sam) and a possession (car). drawing on research in linguistics and other cognitive domains, we propose in section 1.3 three hypotheses about the prominence of animate and inanimate possessions: the animacy hypothesis, possessive hypothesis, and interaction hypothesis. we tested these hypotheses using a sentence continuation task, where we analyzed how likely participants were to mention referents from a prompt sentence depending on the referents’ possession status (possessed vs. indefinite) and animacy (human role nouns vs. alienable concrete objects). given that we observed significant interactions of animacy and possession in both the subject-position and entire-continuation analyses, our results support the interaction hypothesis. because participants were more likely to mention possessed animates (both as subjects and in general) in excess of the combined independent effects of animacy and possession, we conclude that possessed animates get an exceptional boost in discourse prominence—beyond what is predicted simply based on the independent effects of animacy and possession. we suggest that possessed animates’ special status in discourse may relate to the fact that they explicitly denote interpersonal relationships. prior work in the diverse research domains of health, social cognition, and memory has shown that an individual’s relationships relative to other humans are of critical importance. we suggest that the exceptional discourse behavior of possessed animates observed in the current study is linked to a more general cognitive privilege for interpersonal relationships. regarding prior work on the discourse behavior of possessives, our results also address a debate within the centering theory literature mentioned in section 1.1. researchers in this field have proposed two competing accounts for the relative prominence of possessors and possessions in english s-genitives. the first is that possessors are always more prominent than possessions (chae, 2003). alternatively, di eugenio (1998) has proposed that the prominence ranking depends on the animacy of the possession. according to this account, possessors are more prominent than inanimate possessions, but animate possessions are more prominent than their possessors. the results of our experiment show that animate and inanimate possessions behave quite differently on the discourse level, with animate possessions receiving more prominence than their animacy alone would predict. therefore, our results are more compatible with di eugenio’s account, which differentiates the discourse behavior of animate and inanimate possessions and confers extra prominence on possessed animates. finally, some readers may wonder why we did not find a clear overall effect of possession— why are possessed nouns not mentioned more (or perhaps less) often than indefinites? it seems reasonable to assume that possession must have some baseline effect on the discourse representation of a referent, irrespective of animacy; however, the reason our results do not show this effect proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 269 https://doi.org/10.3765/elm https://www.elm-conference.net/ may be attributable to our comparison of possessives to indefinites. we anticipate that choosing a different class of nominals to compare with possessives (e.g. definites) might result in a detectable baseline possessive effect. crucially, however, we do not expect that discourse-related differences between definites, indefinites, and possessives (e.g. givenness, specificity) would systematically vary across our animate and inanimate conditions. therefore, these factors cannot explain the animacy-by-possession interactions we observe in the current experiment. furthermore, in future work comparing possessives to definites, we would still predict that animacy would interact with possession, yielding a result that supports the interaction hypothesis. references ariel, mira. 1988. referring and accessibility. journal of linguistics 24. 65-87. https://doi.org/10.1017/s0022226700011567 arnold, jennifer e. 2001. the effect of thematic roles on pronoun use and frequency of reference continuation. discourse processes 31. 137-162. https://doi.org/10.1207/s15326950dp3102_02 arnold, jennifer and griffin, zenzi m. 2007. the effect of additional characters on choice of referring expression: everyone counts. journal of memory and language 56. 521-536. https://doi.org/10.1016/j.jml.2006.09.007 barker, chris. 2000. definite possessives and discourse novelty. theoretical linguistics 26. 211-227. https://doi.org/10.1515/thli.2000.26.3.211 barr, dale j., levy, roger, scheepers, christoph and tily, harry j. 2013. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language 68. 255-278. https://doi.org/10.1016/j.jml.2012.11.001 bates, douglas, mächler, martin, bolker, ben and walker, steve. 2015. fitting linear mixedeffects models using lme4. journal of statistical software 67. 1-48. https://doi.org/10.18637/jss.v067.i01 bock, kathryn, loebell, helga and morey, randal. 1992. from conceptual roles to structural relations: bridging the syntactic cleft. psychological review 99. 150-171. https://doi.org/10.1037//0033-295x.99.1.150 bonin, patrick, gelin, margaux and bugaiska, aurélia. 2014. animates are better remembered than inanimates: further evidence from word and picture stimuli. memory & cognition 42. 370-382. https://doi.org/10.3758/s13421-013-0368-8 branigan, holly p., pickering, martin j. and tanaka, mikihiro. 2008. contributions of animacy to grammatical function assignment and word order during production. lingua 118. 172189. https://doi.org/10.1016/j.lingua.2007.02.003 van den broek, paul, risden, kirsten, fletcher, charles r. and thurlow, richard. 1996. a “landscape” view of reading: fluctuating patterns of activation and the construction of a stable memory representation. in b. k. britton & a. c. graesser (eds.), models of understanding text, 165-187. lawrence erlbaum associates, inc. cacioppo, john t. and hawkley, louise c. 2009. perceived social isolation and cognition. trends in cognitive sciences 13. 447-454. https://doi.org/10.1016/j.tics.2009.06.005 cacioppo, john t., hughes, mary elizabeth, waite, linda j., hawkley, louise c. and thisted, ronald a. 2006. loneliness as a specific risk factor for depressive symptoms. psychology and aging 21. 140-151. https://doi.org/10.1037/0882-7974.21.1.140 proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 270 https://doi.org/10.3765/elm https://www.elm-conference.net/ chae, sook-hee. 2003. possessives in naturally occurring discourse: a centering approach. university of pennsylvania working papers in linguistics 9. https://repository.upenn.edu/pwpl/vol9/iss1/6 chafe, wallace l. 1976. givenness, contrastiveness, definiteness, subjects, topics, and point of view. in charles n. li (ed.), subject and topic, 25-55. new york: academic press. crawley, rosalind a., stevenson, rosemary j. and kleinman, david. 1990. the use of heuristic strategies in the interpretation of pronouns. journal of psycholinguistic research 19. 245264. https://doi.org/10.1007/bf01077259 dahl, östen. 2008. animacy and egophoricity: grammar, ontology and phylogeny. lingua 118. 141-150. https://doi.org/10.1016/j.lingua.2007.02.008 dahl, östen and fraurud, kari. 1996. animacy in grammar and discourse. in thorstein fretheim and jeanette k. gundel (eds.), reference and referent accessibility, 47-64. amsterdam: benjamins. https://doi.org/10.1075/pbns.38.04dah di eugenio, barbara. 1998. centering in italian. in marilyn walker, aravind joshi and ellen prince (eds.), centering theory in discourse, 115-137. https://arxiv.org/abs/cmplg/9608007 eisenberger, naomi i. and cole, steve w. 2012. social neuroscience and health: neurophysiological mechanisms linking social ties with physical health. nature neuroscience 15. 669674. https://doi.org/10.1038/nn.3086 fisher, ronald p. and craik, fergus i. m. 1980. the effects of elaboration on recognition memory. memory & cognition 8. 400-404. https://doi.org/10.3758/bf03211136 fukumura, kumiko and van gompel, roger p. g. 2011. the effect of animacy on the choice of referring expression. language and cognitive processes 26. 1472-1504. https://doi.org/10.1080/01690965.2010.506444 givón, t. 1983. topic continuity in discourse: a quantitative cross-language study. john benjamins publishing company: philadelphia. https://doi.org/10.1075/tsl.3 gordon, peter c., grosz, barbara j. and gilliom, laura a. 1993. pronouns, names, and the centering of attention in discourse. cognitive science 17. 311-347. https://doi.org/10.1207/s15516709cog1703_1 grosz, barbara j., joshi, aravind k. and weinstein, scott. 1995. centering: a framework for modeling the local coherence of discourse. computational linguistics 21. 203-225. https://www.aclweb.org/anthology/j95-2003 gundel, jeanette k., hedberg, nancy and zacharski, ron. 1993. cognitive status and the form of referring expressions in discourse. language 69. 274-307. https://doi.org/10.2307/416535 hartshorne, joshua k. and snedeker, jesse. 2013. verb argument structure predicts implicit causality: the advantages of finer-grained semantics. language and cognitive processes 28. 1474-1508. https://doi.org/10.1080/01690965.2012.689305 haspelmath, martin. 2017. explaining alienability contrasts in adpossessive constructions: predictability vs. iconicity. zeitschrift für sprachwissenschaft 36. 193-231. https://doi.org/10.1515/zfs-2017-0009 hofmeister, philip. 2011. representational complexity and memory retrieval in language comprehension. language and cognitive processes 26. 376-405. https://doi.org/10.1080/01690965.2010.492642 proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 271 https://doi.org/10.3765/elm https://www.elm-conference.net/ holt-lunstad, julianne, smith, timothy b. and layton, j. bradley. 2010. social relationships and mortality risk: a meta-analytic review. plos medicine 7. e1000316. https://doi.org/10.1371/journal.pmed.1000316 kaiser, elsi. 2009. investigating effects of structural and information‐structural factors on pronoun resolution. in malte zimmermann and caroline féry (eds.), information structure: theoretical, typological, and experimental perspectives, 332-353. oxford: oxford university press. https://doi.org/10.1093/acprof:oso/9780199570959.003.0014 karimi, hossein and ferreira, fernanda. 2016. informativity renders a referent more accessible: evidence from eyetracking. psychonomic bulletin & review 23. 507-525. https://doi.org/10.3758/s13423-015-0917-1 kehler, andrew, kertz, laura, rohde, hannah and elman, jeffrey l. 2008. coherence and coreference revisited. journal of semantics 25. 1-44. https://doi.org/10.1093/jos/ffm018 kehler, andrew and rohde, hannah. 2013. a probabilistic reconciliation of coherence-driven and centering-driven theories of pronoun interpretation. theoretical linguistics 39. 1-37. https://doi.org/10.1515/tl-2013-0001 miller, gregory, chen, edith and cole, steve w. 2009. health psychology: developing biologically plausible models linking the social world and physical health. annual review of psychology 60. 501-524. https://doi.org/10.1146/annurev.psych.60.110707.163551 nairne, james s. 2010. adaptive memory: evolutionary constraints on remembering. in brian h. ross (ed.), psychology of learning and motivation, 1-32: academic press. https://doi.org/10.1016/s0079-7421(10)53001-9 nairne, james s., vanarsdall, joshua e., pandeirada, josefa n. s., cogdill, mindi and lebreton, james m. 2013. adaptive memory: the mnemonic value of animacy. psychological science 24. 2099-2105. https://doi.org/10.1177/0956797613480803 prat-sala, mercè and branigan, holly p. 2000. discourse constraints on syntactic processing in language production: a cross-linguistic study in english and spanish. journal of memory and language 42. 168-182. https://doi.org/10.1006/jmla.1999.2668 rosenbach, anette. 2002. genitive variation in english. de gruyter: berlin/boston. https://doi.org/10.1515/9783110899818 stevenson, rosemary j., crawley, rosalind a. and kleinman, david. 1994. thematic roles, focus and the representation of events. language and cognitive processes 9. 519-548. https://doi.org/10.1080/01690969408402130 tilvis, r. s., kahonen-vare, m. h., jolkkonen, j., valvanne, j., pitkala, k. h. and strandberg, t. e. 2004. predictors of cognitive decline and mortality of aged people over a 10-year period. the journals of gerontology: series a 59. m268-274. https://doi.org/10.1093/gerona/59.3.m268 troyer, melissa, hofmeister, philip and kutas, marta. 2016. elaboration over a discourse facilitates retrieval in sentence processing. frontiers in psychology 7. 374. https://doi.org/10.3389/fpsyg.2016.00374 vanarsdall, joshua e., nairne, james s., pandeirada, josefa n. s. and blunt, janell r. 2012. adaptive memory: animacy processing produces mnemonic advantages. experimental psychology 60. 172-178. https://doi.org/10.1027/1618-3169/a000186 walker, marilyn a. and prince, ellen f. 1996. a bilateral approach to givenness: a hearerstatus algorithm and a centering algorithm. in thorstein fretheim and jeanette k. gundel (eds.), reference and referent accessibility, 291-312. amsterdam: benjamins. https://doi.org/10.1075/pbns.38.17wal proceedings of elm 1: 261-272, 2021 jesse storbeck and elsi kaiser: discourse behavior of possessives reflects the importance of interpersonal relationships. 272 https://doi.org/10.3765/elm https://www.elm-conference.net/ coming in, or going out? measuring the effect of discourse factors on perspective prominence carolyn jane anderson* abstract. although perspectival expressions are a diverse group, they share a common property: their meanings depend on the perspective of a discourse-given individual whose identity is under-specified. this paper investigates how perspective holder prominence is determined through a series of forced choice experiments on american english motion verbs exploring a number of discourse factors: definiteness, mention order, topicality, and subjecthood. the results suggest that both global and local prominence effects play a role in determining how perspectival motion verbs are interpreted. keywords. perspective; discourse prominence; psycholinguistics 1. introduction. perspective plays a role in many linguistic phenomena, including predicates of personal taste, epithets, expressives, motion verbs like come, and discourse environments like free indirect discourse. despite the diversity of perspectival expressions, they share a common property: their meanings depend on the perspective of a discourse-given individual whose identity is underspecified. for instance, in (1), tasty may be anchored to the speaker or to the attitude holder. (1) susan says that the tasty pumpkin pie is almost ready to eat. a. speaker-oriented: it was nice of her to make it when she doesn’t even like pie. b. attitude-holder-oriented: i’m not a fan of pie, but she thinks it’s the best part of fall. the choice of perspective rests on the discourse prominence of the characters: certain characters hold perspectives that are more accessible than others. while the factors that determine discourse prominence have been much discussed in the context of pronominalization (chafe 1976, hobbs 1979, kehler 2002, kehler et al. 2008, kehler & rohde 2013), there is comparatively little work on what affects the prominence of various potential perspective holders in a discourse. this paper contributes to the growing body of work on perspective prominence (harris 2012, kaiser & lee 2017a,b, hinterwimmer 2019) by providing experimental data on a relatively underexplored class of perspectival expressions: motion verbs like come and go. i present a series of experiments that explore how discourse factors impact the prominence of perspective holders for come in english. in experiment 1, i find that both subjecthood and topicality affect perspective prominence. experiments 2 and 3 replicate and delve into these findings: experiment 2 compares two kinds of topicality, and experiment 3 compares two kinds of subjecthood. the results suggest that global effects like topicality and local effects like subjecthood both play a role in determining who the perspective holder is for perspectival motion verbs in english. 2. perspectival expressions. perspectival expressions are interpreted relative to the point-of-view of a particular individual, referred to as the perspective holder.1 for instance, the motion verb *many thanks to daniel altshuler, brian dillon, elsi kaiser, carina bolaños lewen, and the participants and reviewers of experiments in linguistic meaning 1 for their questions and suggestions. authors: carolyn jane anderson, university of massachusetts, amherst (carolynander@umass.edu). 1in the literature on some perspectival expressions, the term judge is used; i use perspective holder for consistency. proceedings of elm 1: 001-014, 2021 c©2021 carolyn jane anderson published by the lsa with permission of the author(s) under a cc by license. 1 https://doi.org/10.3765/elm https://www.elm-conference.net/ come requires a perspective holder to be located at the destination of motion (2).2 (2) context: susie and suzy are chatting in boston. a. suzy: taylor swift is {coming/#going} to boston. b. suzy: taylor swift is {#coming/going} to l.a. the perspective holder plays a critical role in determining the truth conditions of perspectival expressions. a number of semantic treatments have been proposed, which fall into three broad categories: indexical approaches (oshima 2006a,b, korotkova 2016, sudo 2018); binding approaches, which treat the perspective holder as a variable bound by an antecedent (pearson 2013, nishigauchi 2014, charnavel 2018, sundaresan 2018, charnavel 2020); and anaphoric approaches (barlew 2017, roberts 2020). figure 1 shows the semantics of come in each kind of analysis. indexical analysis (oshima 2006b): [[come]]c,g = λx.∃e.∃y ∈ cps.move(e) ∧ dest(e, x) ∧ x = loc(y)), where cps is context parameter c’s set of contextually relevant perspectives. logophoric binding analysis (charnavel 2020): [[comei]]c,g = λx.∃e.move(e)∧dest(e, x)∧ x = loc(li), where li is bound by the subject of a logophoric operator oplog. anaphoric analysis (barlew 2017): [[come]]c,g = λx.∃e.move(e)∧dest(e, x)∧x = loc(p), where p is a prominent perspective holder in the common ground. figure 1: three semantic analyses for come although these approaches propose different ways of valuing the perspectival variable, each assumes a role for pragmatics. in oshima (2006b)’s approach, the context parameter contains a contextually determined set of perspective holders; in charnavel (2020)’s approach, the perspectival variable is bound by a logophoric pronoun whose value is contextually determined; and in barlew (2017)’s approach, the perspectival variable is resolved by the discourse context. the issue of perspective prominence is therefore important regardless of the analysis that is adopted. 3. perspective prominence. the perspective holder can’t be just any individual. this is shown in (2-b): although l.a. is full of people, they cannot serve as perspective holders for come. to be a perspective holder, an individual must hold a discourse prominent perspective: (3) discourse prominence (heusinger & schumacher 2019): a relational property that (1) singles out members of a discourse-given set, (2) shifts as the discourse unfolds, and (3) licenses special access to, or more operations on, prominent elements than competitors. for perspectival expressions, the idea is that there is a ranked set of discourse-given perspectives (or individuals with perspectives), and higher ranked perspectives are more accessible. this accounts for the contrast in (2): in (2-a), the speaker and listener hold discourse-important perspectives in boston, but in (2-b), there is no perspective in l.a. prominent enough to license come. 2note that the set of licit perspective holders varies by language; english is a useful language for exploring perspective prominence because it allows a large set. see nakazawa (2007) and barlew (2017) for more discussion. proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 2 https://doi.org/10.3765/elm https://www.elm-conference.net/ in conversational contexts like (2), it is intuitive that the speaker and listener perspectives are prominent, since they are always relevant to the discussion. but how is discourse prominence determined in third-person narratives, where there may not be any overt conversation participants? one idea is that prominence is determined by how closely related each individual is to the topic of the discourse. this is a global notion of prominence: certain individuals are more important to the discourse theme and have prominence throughout the discourse.3 another possibility is that prominence is determined by local effects: features of the grammatical environment in which the perspectival expression occurs. for example, filling certain thematic or syntactic roles may temporarily increase the prominence of an individual’s perspective. the global and local views of discourse prominence are not necessarily in opposition, since it is also possible that multiple factors affect perspective prominence. 3.1. previous work on perspective prominence. both global and local factors have been explored in the growing body of work on perspective prominence. early work sought to evaluate the relative prevalence of subjectand attitude holder-oriented readings (harris & potts 2009, harris 2012, kaiser 2015). attitude reports are one common environment for global and local prominence to conflict, because speaker perspectives are globally prominent, while attitude holder perspectives are locally prominent in the environment of the attitude report. more recent work has begun to investigate perspective prominence in third-person narrative contexts, focusing on two main perspectival phenomena: free indirect discourse and predicates of personal taste. kaiser has explored the factors that affect how predicates of personal taste are interpreted in a series of perspective-judgment studies (kaiser & lee 2017a,b, kaiser 2020). her results suggest that individuals who fill an experiencer thematic role (kaiser & lee 2017b,a) or who are the source of information (i.e. subject of told or object of heard from) (kaiser 2020) are preferred perspective holders for predicates of personal taste. these are both local prominence effects. however, she interprets her findings as support for an experiencer argument in the semantics of predicates of personal taste (bylinina 2014, mcnally & stojanovic 2017); if this is the case, they may not generalize to perspectival expressions that do not have a grammatically represented experiencer. for free indirect discourse,4 both local and global effects seem to play a role. hinterwimmer (2019) proposes that two factors affect the accessibility of perspectives for free indirect discourse: topicality and thematic roles. he finds that individuals who fill an experiencer thematic role in a preceding sentence, regardless of whether they are the subject, are more likely to be interpreted as free indirect discourse perspective holders. he also finds that individuals introduced early in a discourse are more likely to be perspective holders, evidence of global prominence. meuser et al. (2020) explored these factors experimentally in a series of acceptability judgment and eye-tracking studies. although the eye-tracking results did not show any reliable effects, there was a numerical preference for the perspectives of subjects and of named characters. to summarize, the existing literature on perspectival expressions suggests that both global and local prominence play a role in determining the perspective holder in third-person contexts. on the 3global prominence has been developed in the context of centering theory (chafe 1976), which posits a mechanism for tracking the single most prominent individual (center) at a given point in the discourse and proposes that shifting this center decreases discourse coherence. see poesio et al. (2004) for an overview of different variations. 4a mixed-perspective discourse environment. see banfield (1982), eckardt (2014) for a description. proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 3 https://doi.org/10.3765/elm https://www.elm-conference.net/ one hand, features of the local grammatical context like subjecthood and thematic roles affect the interpretation of free indirect discourse and predicates of personal taste. on the other hand, the work on free indirect discourse suggests that global prominence is also important, since the order and way in which the individuals were introduced affected their prominence. although these findings may not generalize to other perspectival expressions, i take them as useful hints about the discourse factors that may determine perspective prominence. in the following sections, i present a series of experiments that explore the impact of these discourse factors on the interpretation of another class of perspectival expressions: perspectival motion verbs. 4. experiment 1. experiment 1 is a come/go forced choice task measuring the impact of three potential determinants of perspective prominence: subjecthood, definiteness, and mention order.5 the forced choice paradigm makes use of the fact that the subject of come generally cannot be its perspective holder, since a person cannot be at a destination and in motion simultaneously (barlew 2017).6 in a discourse with two characters, we can therefore measure the prominence of one character’s perspective by using the other character as the subject of the motion verb. if come is used, this indicates that the non-subject character’s perspective is prominent; otherwise, go will be used. thus, the rate of come versus go selection is a measure of perspective prominence. 4.1. method. the impact of discourse factors on perspective prominence was measured in a come/go forced choice task. participants read three-sentence narratives and selected come or go for a missing word in the last sentence. figure 2 shows an example item. three discourse factors were manipulated: mention order, definiteness, and subjecthood. 1 definite: larry purchased a sofa for the set of the new play. indefinite: a stage manager purchased a sofa for the set of the new play. 2 a subject: he got a 10% discount on it from a salesclerk. b subject: a salesclerk gave him a 10% discount on it. 3 a perspective: when he {came/went} to deliver it, he got free tickets for the show. b perspective: when he {came/went} to pick it up, he gave him free tickets for the show. figure 2: experiment 1 example item mention order the impact of mention order was measured by comparing the come selection rate for person a and b. in each narrative, person a was introduced in the first sentence and person b was introduced in the second sentence. definiteness the impact of definiteness was measured by manipulating whether or not person a was introduced by name in the first sentence. person b was not named in either condition. subjecthood the impact of subjecthood was measured by manipulating the argument structure of the second sentence. in the a subject condition, person a was the subject, while in the b 5all three experiments were preregistered through the open science foundation: https://osf.io/qmc7t. 6there are exceptions: if a past motion event is being described and the speaker is currently located at the destination of motion, their perspective can be used to anchor come. proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 4 https://doi.org/10.3765/elm https://www.elm-conference.net/ subject condition, person b was the subject. 4.1.1. participants. monolingual american english-speaking participants (n = 64) were recruited through prolific. data from participants (n = 4) who failed to achieve 80% on the attention check items or responded incoherently to one of the two bot check questions was excluded. 4.1.2. materials. all items consisted of three-sentence narratives about two characters. participants were asked to choose between two words to fill in the blank in the third sentence. 32 three-sentence narratives were constructed. there were 8 conditions, created by manipulating (1) whether person a, the discourse-first individual, was named; (2) whether person a or b was the subject of the second sentence; and (3) whether the subject of the motion verb was person a or b. the third sentence serves to disambiguate the perspective holder: when its subject is person a, the rate of come selection measures the prominence of person b’s perspective. two kinds of fillers and two kinds of attention check questions were used (figure 3). the attention check items involved negative polarity item licensing and modals; participants who failed to achieve 80% accuracy on them were excluded. easy: one of the students in the class hadn’t turned in his homework. ms. morris frowned and thought about what to do. after school, she called his mother to {tell / fail} her. hard: jenny was getting dressed for a hike. as she was putting them on, the knee of her jeans ripped. she shrugged and threw them {on / out}. npi: susan was throwing a party on thursday. she had invited a dozen people, but the forecast predicted snow. kevin {doubted / thought} that anyone would come. modal: emma was a vetinarian. she heard that one of her neighbors was adopting a cat. emma told him that he {should / would} get the cat micro-chipped. figure 3: example fillers 4.1.3. procedure. participants saw 4 items in each of 8 conditions, presented using a latin square design. participants also saw 24 filler and attention check items, for a total of 56 items. at the end, participants completed a questionnaire with demographic and bot check questions. 4.2. regression analysis. two mixed-effects logistic regression models were fitted to the response data. the pre-registered model included three fixed effects, perspective, subjecthood, and definiteness, and their interactions. a post-hoc model was used to explore the interaction between definiteness and perspective; it included five fixed effects: perspective, subject, perspective*subject, definiteness nested in perspective a (axname), and definiteness nested in perspective b (bxname). sum contrast coding and the maximal random effects structure were used. 4.3. results. all three manipulated factors (being named, being the first-mentioned individual, and being the subject) were predicted to increase the prominence of a perspective holder, resulting in higher rates of come selection (table 1). as figure 4 shows, the differences in rates of come selection between the person a and person b perspective conditions narrows as the discourse conditions shift from favoring person a to person b. however, the three explored discourse factors proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 5 https://doi.org/10.3765/elm https://www.elm-conference.net/ influence the selection of the perspective holder to differing degrees. ●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●●● ●●●●●●●●●●●● ●●●●●●●●●●●● ●●●● ●●●●●●●● ●●●● ●●●●●●●●●●●● b a def. subj. a indef. subj. a def. subj. b indef. subj. b 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 condition p ar tic ip an t m ea n figure 4: experiment 1 means by condition persp. a b come[95% ci] a name ob 62% [53;71] sub 56% [49;62] anon ob 63% [55;71] sub 50% [41;58] b name ob 21% [16;26] sub 26% [20;32] anon ob 23% [17;30] sub 23% [17;30] table 1: experiment 1 mean rate of come selection the effect of mention order was strongest: in all conditions, the mean rate of come responses was higher for person a’s perspective than person b’s. as table 2 shows, the effect of mention order was significant in both models, and switching perspectives from person b to person a was estimated to increase the rate of come responses to a greater extent than subjecthood or definiteness. fixed effect (n = 2048) β̂ z p (intercept) -0.59(+/0.17) -3.5 <0.001 a perspective 1.02(+/0.14) 7.1 <0.00001 a subject 0.10(+/0.06) 1.8 0.07 a definiteness 0.04(+/0.06) 0.7 0.50 perspective*subject 0.19(+/0.06) 3.4 <0.001 perspective*definiteness 0.04(+/0.06) 0.7 0.51 (intercept) -0.59(+/0.17) -3.5 0.0004 a perspective 1.0(+/0.14) 7.1 <0.00001 definiteness 0.04(+/0.06) 0.68 0.50 axsubject 0.29(+/0.07) 3.9 0.0001 bxsubject -0.09(+/0.08) -1.1 0.29 perspective*definiteness 0.04(+/0.06) 0.66 0.51 table 2: experiment 1 regression analyses, interaction (top) and nested (bottom) models there was also a reliable interaction between subjecthood and perspective (p < 0.001). the nested mixed-effects model found that person a subjecthood had a significant positive effect on come response rates for person a’s perspective (p < 0.001), but did not find a significant effect on come response rates for person b’s perspective. neither the main effect of definiteness nor its interaction with perspective was significant, although a numerical trend is observable in figure 4. 4.4. discussion. the results of experiment 1 support a role for both local and global prominence in determining the perspective holder of come. the strongest observed effect was that of mention order: person a’s perspective was more prominent than person b’s in all conditions. proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 6 https://doi.org/10.3765/elm https://www.elm-conference.net/ mention order is a correlate of global prominence: a character who is mentioned early is more likely to be the topic of the discourse. a local prominence effect was also observed: subjecthood had a reliable, though weaker effect on perspective prominence. although the results of experiment 1 suggest that both global and local prominence are in play, it leaves open several questions about the source of these effects. unlike paradigms from previous work, the experiment 1 paradigm did not distinguish between argument structure and thematic roles, which means that the observed subjecthood effect could also arise from a correlation between thematic roles and subjecthood. the observed effect of mention order might also have several causes: is person a’s perspective more prominent because they are mentioned earlier, or because by appearing early, they are more closely associated with the discourse theme? these factors were conflated in experiment 1: in the figure 2 stimulus, for instance, the stage manager character is both mentioned first and more closely related to the discourse topic. the design in experiment 1 also had a potential pronominalization confound. because prominence drives pronominalization, a character who is referred to by a pronoun might be inferred to be more prominent. in experiment 1, person a was referred to with a pronoun in the second sentence to improve the naturalness of the narratives. however, this may have inadvertently increased the prominence of that character for reasons unrelated to mention order. experiment 2 seeks to address these last two points with an updated experimental paradigm. 5. experiment 2. experiment 2 explores the effect of subjecthood and two indicators of topicality: mention order and relation to the theme of the discourse. similar to experiment 1, the prominence of characters’ perspectives was measured in a come/go forced choice task. 5.1. method. as in experiment 1, participants read three-sentence narratives and selected come or go for a missing word in the last sentence. three discourse factors were manipulated: mention order, relation to the theme of the discourse, and subjecthood. figure 5 shows an example item. 1 mention: freshman drew was struggling with his computer science homework. theme: computer science is one of the most difficult subjects for freshmen. 2 a subject: drew asked an older student sandra for help with the class. b subject: an older student sandra offered to tutor drew. 3 a perspective: when she {came/went} to his dorm, they made a plan for how to study. b perspective: when he {came/went} to her dorm, they made a plan for how to study. figure 5: experiment 2 example item relation to theme as in experiment 1, each narrative involved two characters, a and b. in this experiment, the first sentence set up a topic for the discourse more closely associated with a. for instance, in figure 5, the first sentence frames the narrative around the challenges freshmen face. mention order the impact of mention order was measured by manipulating whether or not person a was introduced in the first sentence. in the mention condition, person a was mentioned in the first sentence, while in the theme condition, neither character was. proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 7 https://doi.org/10.3765/elm https://www.elm-conference.net/ subjecthood the impact of subjecthood was measured by manipulating the argument structure of the second sentence. in the a subject condition, person a was the subject, while in the b subject condition, person b was. the other character was an object of the verb or a prepositional phrase. 5.1.1. participants. monolingual american english-speaking participants (n = 80) were recruited through prolific. the data from participants (n = 15) who failed to achieve 80% on the attention check items or responded incoherently to a bot check question was excluded. 5.1.2. materials. as in experiment 1, items consisted of three-sentence narratives about two characters. 32 narratives were constructed. there were 8 conditions, created by manipulating (1) whether person a was introduced in the first sentence; (2) whether the subject of the second sentence was person a or b; and (3) whether the subject of the motion verb was person a or b. unlike in experiment 1, both characters were mentioned by name in all narratives. in addition, the use of pronouns in the second sentence was avoided. as before, the prominence of a character’s perspective was measured as the rate of come selection. 5.1.3. procedure. participants saw 4 items in each of 8 conditions, presented using a latin square design. participants also saw the 24 filler and attention check items from experiment 1, for a total of 56 items. at the end, participants answered demographic and bot check questions. 5.2. regression analysis. two mixed-effects logistic regression models were fitted to the response data. the pre-registered model included three fixed effects: perspective, subjecthood, and theme, and their interactions. a post-hoc nested model was used to explore the interaction effects. it included five fixed effects: perspective, subjecthood nested in perspective a, subjecthood nested in perspective b, theme nested in perspective a, and theme nested in perspective b. sum contrast coding and the maximal random effects structure were used. 5.3. results. the results of experiment 2 are shown in table 3. they support a role for all three discourse factors explored. as figure 6 shows, as subjecthood and mention order are manipulated, the rates of come selection increase for perspective b and decrease for perspective a. ●●●● b a mention, subj. a theme, subj. a mention, subj. b theme, subj. b 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 condition p ar tic ip an t m ea n figure 6: experiment 2 means by condition persp. a b come[95% ci] a mention ob 45% [41;48] sub 57% [54;60] theme ob 32% [29;35] sub 52% [50;55] b mention ob 25% [23;38] sub 37% [35;40] theme ob 26% [23;28] sub 49% [46;52] table 3: experiment 2 mean rate of come selection as in experiment 1, there was a general preference for perspective a, indicating that being proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 8 https://doi.org/10.3765/elm https://www.elm-conference.net/ related to the theme of the discourse increased the prominence of person a’s perspective. as table 4 shows, the observed effect of perspective was significant in the mixed-effects model (p < 0.001). fixed effect (n = 2560) β̂ z p (intercept) -0.48(+/0.13) -3.5 0.0004 a perspective 0.3(+/0.08) 3.8 0.0001 a subject -0.02(+/0.05) -0.54 0.59 theme 0.03(+/0.05) 0.52 0.60 a perspective*subject 0.43(+/0.05) 9.5 <0.00001 a perspective*theme 0.18(+/0.05) 4.1 <0.00001 (intercept) -0.51(+/0.14) -3.6 0.0004 a perspective 0.33(+/0.08) 3.9 <0.00001 axa subject 0.42(+/0.08) 5.0 <0.00001 bxa subject -0.48(+/0.1) -5.0 <0.00001 axtheme 0.22(+/0.07) 3.1 0.002 bxtheme -0.17(+/0.08) -2.0 0.04 table 4: experiment 2 regression analyses, interaction (top) and nested (bottom) models a strong effect of subjecthood was also observed: within each mention/theme condition, come selection increased when the perspective holder was the subject of the second sentence. in the theme b subject condition, unlike all other conditions, the mean come rate was higher for person b’s perspective than a’s. the interaction of perspective and subject was significant (p < 0.00001). in the nested model (table 4), both nested effects of subject were significant (p < 0.00001), but in opposite directions: person a having subjecthood increased the rate of come selection when a was the perspective holder, but decreased it when b was the perspective holder. the results also suggest that mention order has an effect beyond thematic relation: come selection rates were higher for perspective a when person a was introduced in the first sentence. the interaction between perspective and theme/mention was significant (p < 0.00001), and in the nested model, there were significant effects of theme for both levels of perspective, again in different directions: introducing person a in the first sentence increased the rate of come selections for perspective a (p < 0.01), while decreasing it for perspective b (p < 0.05). however, the effect for perspective b was weak when b was not the second sentence’s subject. 5.4. discussion. the results of experiment 2 support the hypothesis that both global and local prominence factors affect perspective holder selection for come. this is seen most clearly in the theme b subject condition. although person a’s perspective was generally preferred due to its global prominence, when this was weakened by removing the mention order effect, the local prominence granted to person b by its subjecthood in the second sentence was enough to overcome a’s global prominence and produce higher rates of come selection for b’s perspective. in addition, the data provide a more nuanced look into the global prominence effect: both mention order and relation to the theme of the discourse increase perspective prominence. in reality, the effect of relation to the theme may be stronger than what was observed: some theme narratives may have failed to associate only person a with the theme, since it was challenging to do this and preserve discourse coherence. proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 9 https://doi.org/10.3765/elm https://www.elm-conference.net/ 6. experiment 3. experiment 3 explores the subjecthood effect observed in the previous two experiments by comparing two kinds of subject environments. since hinterwimmer (2019) posits that sentience with respect to a preceding eventuality influences prominence, we might expect subjects of attitude verbs to receive even larger boosts in prominence than other subjects.7 experiment 3 explores this hypothesis by comparing the prominence of subjects of attitude verbs with subjects of verbs with agent/patient thematic roles. 6.1. method. as before, the prominence of characters’ perspectives was measured in a come/go forced choice task. participants read three-sentence narratives and selected a missing word in the last sentence. three discourse factors were manipulated: relation of the character to the theme of the discourse, the subject of the second sentence, and whether the second sentence contained an attitude verb or an verb with agent/patient thematic roles. figure 7 shows an example item. 1 being a bus driver requires a lot of patience since children can be careless. 2 a attitude subject: sharon was annoyed to find fifth-grader nick’s lunch on one of the seats in her bus. a a/p subject: sharon found fifth-grader nick’s lunch on one of the seats in her bus. b attitude subject: fifth-grader nick regretted leaving his lunch on one of the seats in sharon’s bus. b a/p subject: fifth-grader nick had left his lunch on one of the seats in sharon’s bus. 3 a perspective: when he {came/went} to pick it up, he said he was sorry. b perspective: when she {came/went} to pick him up, she scolded him. figure 7: experiment 3 example item relation to theme as in the experiment 2 theme condition, each narrative began with a thematic first sentence that was more closely related to person a than person b. subjecthood as in the previous two experiments, the impact of subjecthood was measured by manipulating whether person a or person b was the subject of the second sentence. attitude holder the impact of being an attitude holder was manipulated within the second sentence as well. in the a/p condition, the matrix verb was one with agent/patient thematic roles. in the attitude condition, the matrix verb was an attitude predicate. the subject of the embedded verb was always the same as the subject of the attitude verb. 6.1.1. participants. monolingual american english-speaking participants (n = 80) were recruited through prolific. the data from participants (n = 9) who failed to achieve 80% on the attention check items or who gave an incoherent response to one of the two free response bot check questions in the debriefing questionnaire was excluded. 6.1.2. materials. the theme condition stimuli from experiment 2 were adapted for this experiment. 32 three-sentence narratives were constructed. there were 8 conditions, created by 7roberts (2020) proposes that attitude verbs introduce attitude holder perspectives into the discourse context. proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 10 https://doi.org/10.3765/elm https://www.elm-conference.net/ manipulating (1) the subject of the second sentence; (2) whether the second sentence contained an attitude verb or an agent/patient verb; and (3) the subject of the motion verb. 6.1.3. procedure. participants saw 4 items in each of 8 conditions, presented using a latin square design. participants also saw the 24 filler and attention check items from experiment 1, for a total of 56 items. afterwards, participants answered demographic and bot check questions. 6.2. regression analysis. two pre-registered mixed-effects logistic regression models were fitted to the response data. the first included three fixed effects, perspective, subjecthood, and attitude, and their interactions, while the second treated subjecthood and attitude as nested within perspective. both used sum contrast coding and the maximal random effects structure. 6.3. results. experiment 3 replicated the subjecthood effect observed in the prior two experiments, but did not find a reliable effect of topicality or of type of subjecthood. figure 8 shows the participant means by condition. within each of the four narrative conditions, the rate of come selection is consistently higher for the character who is the subject of the second sentence. this supports the subjecthood effect observed in the previous experiments. as table 6 shows, the interaction between subjecthood and perspective was significant (p < 0.00001), as were the nested effects of subjecthood for both person a and b (p < 0.00001 each). ●●●●●●●●●● ●●●●●● ●●●●●●●●●●●●●●● b a att. holder a agent a agent b att. holder b 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 condition p ar tic ip an t m ea n figure 8: experiment 3 means by condition persp. a b come[95% ci] a ap ob 33% [30;35] sub 47% [44;50] attitude ob 30% [28;33] sub 49% [47;51] b a/p ob 27% [24;29] sub 45% [43;47] attitude ob 27% [25;29] sub 44% [42;47] table 5: experiment 3 mean rate of come selection however, the effect of relation to the theme observed in experiment 2 did not replicate. although the means for person a conditions were higher in each pair of equivalent conditions (for instance, comparing the person a agent means with the person b agent means), the difference is slight, and as table 5 shows, the confidence intervals largely overlap. the main effect of perspective was not significant in the mixed-effects model (p = 0.1). manipulating the syntactic environment in the second sentence did not have a reliable effect: while the rate of come selection was slightly higher in the attitude verb condition when person a was the subject and perspective holder, it was lower when person b was the subject and perspective holder. no significant effect was found in either mixed-effects model (table 6). 6.4. discussion. the results of experiment 3 partially replicate the findings of the previous two experiments: they support an effect of subjecthood on perspective prominence, but not of relation proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 11 https://doi.org/10.3765/elm https://www.elm-conference.net/ fixed effect (n = 3200) β̂ z p (intercept) -0.63(+/0.12) -5.1 <0.00001 a perspective 0.1(+/0.07) 1.6 0.1 a subject -0.03(+/0.05) -0.6 0.55 attitude -0.003(+/0.04) -0.06 0.9 a perspective*subject 0.44(+/0.05) 9.3 <0.00001 a perspective*attitudes -0.005(+/0.04) -0.1 0.9 (intercept) -0.64(+/0.13) -5.1 <0.00001 a perspective 0.1(+/0.07) 1.5 0.14 axa subject 0.43(+/0.08) 5.2 <0.00001 bxa subject -0.47(+/0.07) -6.3 <0.00001 axattitude -0.007(+/0.06) -0.1 0.91 bxattitude 0.002(+/0.06) 0.04 0.97 table 6: experiment 3 regression analyses, interaction (top) and nested (bottom) models to the theme. they also do not provide insight into the effect of subjecthood: although attitude reports were predicted to strengthen the effect of subjecthood, no reliable effect was observed. 7. conclusion. i have presented three experiments exploring the discourse factors that affect the prominence of perspective holders for perspectival motion verbs in american english. the results suggest that both local and global prominence play a role. experiment 1 provided evidence that topicality and subjecthood increase the accessibility of characters’ perspectives for come. both of these effects were replicated in subsequent experiments. however, the experiments exploring the source of these effects were less successful. while experiment 2 found two topicality effects, mention order and relation to theme, the latter did not replicate in experiment 3. experiment 3 also explored two kinds of subject environments, attitude verbs and agent/patient verbs, but found no reliable difference between them. the experimental data presented here contribute to the growing body of work exploring how perspective prominence is determined. it supports a view that although certain characters have global prominence by virtue of their topicality, others can become temporarily more prominent because of features of the syntactic environments in which they appear. although these findings are consistent with those from work on other perspectival expressions (kaiser & lee 2017a, hinterwimmer 2019), there remain many open questions about how consistent the observed effects are across classes of perspectival expressions. because perspective prominence is a discourse-level phenomena, we might expect the same perspective ranking to be utilized across classes of expressions. on the other hand, if some of the attested effects spring from the semantics of individual expressions, as kaiser & lee (2017a) suggest, then perspective prominence might not generalize across classes of perspective-sensitivity. references banfield, ann. 1982. unspeakable sentences: narration and representation in the language of fiction. london: routledge. https://doi.org/10.4324/9781315746609. proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 12 https://doi.org/10.3765/elm https://www.elm-conference.net/ barlew, jefferson. 2017. the semantics and pragmatics of perspectival expressions in english and bulu: the case of deictic motion verbs. columbus, ohio: the ohio state university dissertation. bylinina, lisa. 2014. the grammar of standards. utrecht: lot dissertation series. chafe, wallace. 1976. givenness, contrastiveness, definiteness, subjects, and topics. in charles li (ed.), subject and topic, 25–76. new york: academic press. charnavel, isabelle. 2018. deictic perspective and logophoric exemption from condition a. in william g. bennett, lindsay hracs & dennis ryan storoshenko (eds.), west coast conference on formal linguistics, vol. 35, 124–131. somerville, ma: cascadilla proceedings project. charnavel, isabelle. 2020. logophoricity and locality: a view from french anaphors. linguistic inquiry 51(4). 671–723. eckardt, regine. 2014. the semantics of free indirect discourse: how texts allow us to mindread and eavesdrop (current research in the semantics / pragmatics interface 31). boston: brill. https://doi.org/10.1163/9789004266735. harris, jesse a. 2012. processing perspectives. amherst, ma: university of massachusetts, amherst dissertation. https://scholarworks.umass.edu/dissertations/ aai3498347. harris, jesse a. & christopher potts. 2009. perspective-shifting with appositives and expressives. linguistics and philosophy 36(2). 523–552. https://doi.org/10.1007/s10988-010-9070-5. heusinger, klaus von & petra b. schumacher. 2019. discourse prominence: definition and application. journal of pragmatics 154. 117–127. https://doi.org/10.1016/j. pragma.2019.07.025. hinterwimmer, stefan. 2019. prominent protagonists. journal of pragmatics 154. 79–91. https://doi.org/10.1016/j.pragma.2017.12.003. hobbs, jerry r. 1979. coherence and coreference. cognitive science 3(1). 67–90. https://doi.org/10.1207/s15516709cog03014. kaiser, elsi. 2015. perspective-shifting and free indirect discourse: experimental investigations. in sarah d’antonio, mary moroney & carol rose little (eds.), semantics and linguistic theory, vol. 25, 346–372. https://doi.org/10.3765/salt.v25i0.3436. kaiser, elsi. 2020. shifty behavior: investigating predicates of personal taste and perspectival anaphora. in joseph rhyne, kaelyn lamp, nicole dreier & nicole kwon (eds.), semantics and linguistic theory, vol. 30, in press. https://doi.org/10.3765/salt.v30i0.4850. kaiser, elsi & jamie herron lee. 2017a. experience matters: a psycholinguistic investigation of predicates of personal taste. in dan burgdorf, jacob collard, sireemas maspong & brynhildur stefánsdóttir (eds.), semantics and linguistic theory, vol. 27, 323–339. https://doi.org/10.3765/salt.v27i0.4151. kaiser, elsi & jamie herron lee. 2017b. predicates of personal taste and multidimensional adjectives: an experimental investigation. in william g. bennett, lindsay hracs & dennis ryan storoshenko (eds.), west coast conference on formal linguistics, vol. 35, 224–231. somerville, ma: cascadilla proceedings project. kehler, andrew. 2002. coherence, reference, and the theory of grammar. stanford: csli publiproceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 13 https://doi.org/10.3765/elm https://www.elm-conference.net/ cations. kehler, andrew, laura kertz, hannah rohde & jeffrey l. elman. 2008. coherence and coreference revisited. journal of semantics 25(1). 1–44. https://doi.org/10.1093/jos/ffm018. kehler, andrew & hannah rohde. 2013. a probabilistic reconciliation of coherence-driven and centering-driven theories of pronoun interpretation. theoretical linguistics 39(1-2). 1–37. https://doi.org/10.1515/tl-2013-0001. korotkova, natalia. 2016. heterogeneity and uniformity in the evidential domain. los angeles, ca: ucla dissertation. https://escholarship.org/uc/item/40m5f2f1. mcnally, louise & isidora stojanovic. 2017. aesthetic adjectives. in james young (ed.), the semantics of aesthetic judgments, 17–37. oxford: oxford university press. 10.1093/ acprof:oso/9780198714590.001.0001. meuser, sara, stefan hinterwimmer & maximilian hörl. 2020. online processing of protagonists’ perspective-taking. in the cuny sentence processing conference, vol. 33, online. nakazawa, tsuneko. 2007. a typology of the ground of deictic motion verbs as path-conflating verbs: the speaker, addressee, and beyond. poznan studies in contemporary linguistics 43(2). https://doi.org/10.2478/v10010-007-0014-3. nishigauchi, taisuke. 2014. reflexive binding: awareness and empathy from a syntactic point of view. journal of east asian linguistics 23. 157–206. https://doi.org/10.1007/s10831-0139110-6. oshima, david. 2006a. go and come revisited: what serves as a reference point? in zhenya antić, charles b. chang, emily cibelli, jisup hong, michael j. houser, clare s. sandy, maziar toosarvandani & yao yao (eds.), proceedings of the berkeley linguistics society, vol. 32, 287–298. berkeley linguistics society. https://doi.org/10.3765/bls.v32i1.3466. oshima, david. 2006b. motion deixis, indexicality, and presupposition. in masayuki gibson & jonathan howell (eds.), proceedings of semantics and linguistic theory, vol. 16, 172–189. https://doi.org/10.3765/salt.v16i0.2942. pearson, hazel. 2013. a judge-free semantics for predicates of personal taste. journal of semantics 30. 103–154. https://doi.org/10.1093/jos/ffs001. poesio, massimo, rosemary stevenson, barbara di eugenio & janet hitzeman. 2004. centering: a parametric theory and its instantiations. computational linguistics 30(3). 309–363. 10.1162/0891201041850911. roberts, craige. 2020. character study: a de se semantics for indexicals. ms. new york university. sudo, yasutada. 2018. come vs. go and perspective shift. ms. rutgers university. sundaresan, sandhya. 2018. perspective is syntactic: evidence from anaphora. glossa 3(1)(128). 1–40. http://doi.org/10.5334/gjgl.81. proceedings of elm 1: 001-014, 2021 carolyn jane anderson: coming in, or going out? measuring the effect of discourse factors on perspective prominence. 14 https://doi.org/10.3765/elm https://www.elm-conference.net/ pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge fereshteh modarresi and manfred krifka* abstract. there are different theories about the nature of pseudo-incorporated nouns (pins), which feature a non-specific, number-neutral interpretation. for a proper analysis it is crucial to take their anaphoric potential into account. this paper investigates if and how pins introduce discourse referents, with evidence from persian, and which theory matches this behavior best. we report on experiments in which the stereotypical enrichment of the number-neutral interpretation was systematically varied with two types of biases — towards a singular or a plural interpretation — and in the neutral case, when such a bias is lacking. the results of the experiments are compatible with krifka & modarresi (2016), which considers pin objects as dependent singular definites (similar to weak definites) within existential closure over an event variable. keywords. semantics; pragmatics; psycholinguistics; anaphoric reference; weak definites; pseudo-incorporation. 1. pseudo-incorporated nominals (pins). pseudo-incorporated nominals (pins) have been identified by massam (2001) as arguments of verbal predicates that exhibit certain syntactic and semantic properties. as for their syntax, they exhibit certain restrictions with respect to their prosodic and syntactic independence but are not fully morphologically incorporated, as in mary went to church in contrast to mary was a regular churchgoer and mary went to a church. as for their semantics, pins are unspecific and number-neutral, as in the people in the town went to church – no particular church is referred to, in fact there might be more than one church. in their interpretation, pins correspond to weak definites as mary was taken to the hospital, which are expressed by syntactically reduced forms in some languages (e.g., in german ins hospital vs. in das hospital, cf. schwarz (2014)). pins may be realized in different ways in different languages, and they may be more or less prominent in certain languages (see borik & gehrke 2015, massam 2017 and chung & ladusaw 2020 for recent treatments). the current paper deals with direct object arguments in persian, which has a direct object marker in form of a postposition -rā that triggers a specific or definite interpretation with bare nominals, cf. (1). indefinite dps, as marked with the indefinite number word or article yek, can occur with or without rā, cf. (1). bare nominals without -rā are illustrated in (1). (1) a. sara ketāb rā kharid. b. sara yek ketāb rā kharid. sara book om bought sara one book om bought ‘sara bought the book ‘sara bought a book’ * we acknowledge and thank our collaborators fahimeh taheri and hakimeh rezaie for their assistance in data collection. research for this paper was funded by the dfg project dfg kr951/10-1 anapin: anaphoric potential of incorporated nominals and weak definites (anapin) (manfred krifka & werner frey) at leibniz-zentrum allgemeine sprachwissenschaft (zas) berlin. fereshteh modarresi (modarresi@leibniz-zas.de) & manfred krifka (krifka@leibniz-zas.de). abbreviations: pin ‘pseudo-incorporated nominal’, bn ‘bare noun’, yk ‘marked by indefinite yek’, om ‘object marker’, idf: indefinite, ez: ezafe (linker), sg: singular, pl: plural, ø: null, 1: 1st person, 3: 3rd person. proceedings of elm 1: 224-236, 2021 c©2021 fereshteh modarresi and manfred krifka published by the lsa with permission of the author(s) under a cc by license. 224 https://doi.org/10.3765/elm https://www.elm-conference.net/ c. sara yek ketāb kharid. d. sara ketāb kharid. sara one book bought sara one book bought ‘sara bought a book’ ‘sara bought an book / books.’ bare objects without -rā (d) have a number-neutral interpretation; they are also interpreted as nonspecific. as modarresi (2014, 2015) shows, they also have the other defining properties of pins: they form one prosodic domain with the verb, they cannot undergo scrambling but only focus movement, and can be expanded by modifiers. hence they can be analyzed as instances of pins. there are a number of different theories about the nature of pins. as for their semantics, they have been analyzed as referring to kinds (cf. ghomeshi 2008, aguilar-guevara & zwarts 2010, dayal 2011), denoting properties (cf. mcnally & van geenhoven 2005, dobrovie-sorin & giurgea 2015), and involving a restrictive function-argument composition (cf. ladusaw & chung 2003, espinal & mcnally 2011). other accounts have been proposed as well. for a proper analysis of the semantics of pins it is crucial to take their anaphoric potential into account, i.e. their ability to introduce discourse referents (drs) that can be picked up by anaphoric expressions in subsequent discourse. the three theoretical options mentioned above predict that pins are not accessible for anaphoric uptake without further mechanisms, as kinds, properties, and semantic restriction would not result in object-level drs. they predict that pins are “discourse opaque”. there are other theories of the interpretation of pins that predict that anaphoric uptake is possible. modarresi (2014) assumed that pins in persian introduce numberneutral drs, thus capturing their number-neutral interpretation and an assumed preference for uptake by null anaphora, which are number-neutral as well. according to her, pins are “discourse transparent” (cf. also van geenhoven, 1998 for incorporation in greenlandic eskimo). krifka & modarresi (2016) proposed that pins introduce drs with restricted scope that can be extended by an operation of abstraction and summation, allowing for a restricted anaphoric uptake. in a term introduced by farkas & de swart 2003, pins are “discourse translucent” (cf. also yanovich 2008). opinions about anaphoric accessibility of pins have varied widely, with little empirical evidence beyond the intuitions of the researchers or anecdotal observations. but this situation is changing. in a first experimental investigation on the processing of small discourses with bare noun antecedents in mandarin, law & syrett (2017) found evidence for discourse translucency in a self-paced reading experiment. for german, brocher et al. (2020) found evidence for a reduced prominence of pins using eye tracking, in experiments on persian (cf. modarresi & krifka 2020 and to appear a, b) that involved item selections and free sentence completion, we found clear evidence that anaphoric reference to pins is natural, albeit slightly less straightforward than with indefinite pins. we take it that the current experimental evidence speaks against the opacity hypothesis for pins, at least for the investigated types in mandarin, german, and persian. in the present article, we will report on additional experiments to our cited work that bear on the precise mechanism how pins introduce drs. we have seen that pins allow for a number-neutral interpretation, which is in principle compatible with reference to a single object or a multitude of objects, e.g. one or more books in (1). in modarresi & krifka (to appear a, b) we considered experimental items in a way that should prevent object to be understood with a bias towards one or more than one object. this is arguably the case for (1), as sara could equally likely have bought one book or more than one book. this contrasts with (2) and (3), which have a strong tendency towards a multitude vs. a single interpretation, respectively. in (2) this is due to the temporal quantification ‘the whole day’, in (3) this is due to stereotypical knowledge about prizes in competition. proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 225 https://doi.org/10.3765/elm https://www.elm-conference.net/ (2) man tamam-e-ruz livan shost.am va baad (#oon/oona) ro/ø khoshk kardam. i whole-ez-day glass washed-1sg and then (#that/those)-om/ø dry did.1sg i washed glass the whole day and then dried ø / them. (3) nafar-e aval saat-e-tala barandeh shod vali doos-esh/#eshoon/ø na-dasht the first winner watch-ez-gold winner became.3sg but like-it/#them/ø neg-had.3sg ‘the first winner won golden watch but didn’t like it/like’ similarly, (4) tends to be understood as buying a multitude of carrots, whereas (5) is preferably interpreted as involving only one car. the difference is due to the stereotypical enrichment of linguistic information – carrots are typically bought in groups and cars are bought as single object. (4) sara havij kharid va man poost-e-shoon/ ø /?esh ro kandam. sara carrot bought.3sg and i skin-ez-them/ ø /it/ om peeled.1sg. ‘sara bought carrot and i skinned them// ø /?it’ (5) sarah emrooz mashin kharid va rooz-e-baad foroukht/ ø /esh/?eshoon sara today car bought.3sg and day-ez-next sold/∅/it/#them ‘sar bought car today and sold (ø/it) next day’ in the current paper, we will report on experiments in which the stereotypical enrichment of the number-neutral interpretation was systematically varied. this leads to a new evaluation of the precise mechanisms by which pins introduce drs. 2. two theoretical models of anaphoric uptake for pins. as mentioned above, we will concentrate here on theories that are consistent with the finding that anaphora to pins (more specifically, bare object nouns in persian) is possible, following the recent experimental evidence of modarresi & krifka (to appear a, b). there are two ways in which such anaphora may work: directly, by the introduction of drs by antecedent expressions that are picked up by co-referring expressions, or indirectly, by associative anaphora. associative anaphora is illustrated in the following case: (6) sarah bought a book. the cover picture looked interesting. the antecedent clause did not introduce a dr for the cover picture. rather, a referent of this expression can be constructed due to the introduction of a dr for a book, and the stereotypical enrichment that books often have cover pictures on them. the results in the sentence continuation experiment of modarresi & krifka (to appear. a, b) speak against the possibility that anaphora to bare objects in persian is predominantly by associative anaphora, as in this case we should find full dps as the preferred case of anaphoric uptake. as full dps were rarely produced by the participants (see also section 3.3 below), we can exclude them as a dominant way of uptake. hence, we assume that anaphoric uptake of pins in persian is mediated via discourse referents. there are different theoretical options for the uptake mediated by drs, of which we consider two. both theoretical options are couched in the language of discourse representation theory (drt), which we outline here to the extent that is necessary (cf. kamp & reyle 1993, kamp et al. 2011 for comprehensive introductions). drt assumes a semantic representation in terms of discourse representation structures (drss), which are pairs of a set of accessible drs and conditions on these drs. drss are typically depicted in box format but we will render them here more compactly as pairs of the form ⟨discourse referents | conditions⟩. a sentence is interpreted as proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 226 https://doi.org/10.3765/elm https://www.elm-conference.net/ expanding the drs of the preceding discourse, where non-anaphoric dps introduce new drs, and anaphoric dps pick up drs that were already introduced. this can be illustrated for yek-marked as follows: (7) a. sara yek ketāb kharid. b. ø/oo foran khoond-esh. sara one book bought (s)he immediately read-it ‘sarah bought a book.’ ‘she/he immediately read it.’ (8) ⟨ | ⟩ + (5)(a) = ⟨x₁ x₂ e₁ │x₁ = sara, book(x₂), |x₂| = 1, e₁: bought(x₁,x₂)⟩ (9) (13) + (7)(b) = # x₁ x₂ e₁ e₂ % x₁ = sara, book(x₂), |x₂| = 1, e₁: bought(x₁,x₂) immediately_after(e₁,e₂), e₂: read(x₁,x₂) & here, (8) describes the update of an empty initial drs ⟨ | ⟩ by the first sentence, which introduces a dr x₁ for sara, a dr x₂ for one book, and a dr e₁ for a past event of buying of x₂ by x₁. in the conditions, we use the format of kamp & reyle (1993); in particular, |x| specifies the number of atomic entities that the dr x is anchored to, and e: r(x, y) states that e is an event in which x and y stand in the relation r to each other. the resulting drs is further expanded in (9) by the second clause, which picks up x₁ and x₂ by the anaphoric devices of an subject pronoun realized as oo or empty, as persian is a pro-drop language, and an object clitic -esh. the second clause also introduces another event dr e₂ that is immediately after e₁ and is a past reading of x₂ by x₁. drss are interpreted with respect to a model m that contains a set of entities that have certain properties and stand in certain relations to each other. a drs ⟨ d | c ⟩ with a set of drs d and a set of conditions c is true with respect to a model m iff there is a function g that maps all drs in the set d to entities in m such that all conditions in c are true for the corresponding entities in m. the first theory involving drs was proposed by modarresi (2014, 2015). the yek-marked singular indefinite object in (7) introduces a dr that is anchored to a single book (cf. the condition |x₂| = 1). modarresi assumes that bare objects differ minimally insofar as they introduce number neutral drs (already assumed by kamp & reyle 1993 for different phenomena), which are given here by greek letters ξ. this is justified by the number-neutral interpretation of bare objects in persian (and pins in general). (10) a. sara ketāb kharid. b. ø/(oo) khoond-ø/-esh/-eshoon. sara book bought (s)he read-ø/-it/-them ‘sarah bought a book.’ ‘she/he read it.’ (11) ⟨ | ⟩ + (7)(a) = ⟨ x₁ ξ₂ e₁| x₁ = sara, book(ξ₂), |ξ₂| ≥ 1, e₁: bought(x₁,ξ₂) ⟩ (12) (13) + (7)(b) = # ξ₁ x₂ e₁ e₂ % x₁ = sara, book(ξ₂), |ξ₂| ≥ 1, e₁: bought(x₁,ξ₂) e₂: read(x₁,ξ₂), |ξ₂| ≥/=/>1 & we represent the fact that number-neutral drs ξ can be anchored to atomic entities or to sum individuals consisting of more than just one atomic entity by the condition |ξ]≥1. in this case, the anaphoric uptake is natural with a null anaphor, which does not restrict the dr to any particular number. but uptake is also possible by the singular enclitic anaphor -esh and the plural enclitic anaphor -eshoon. in the latter cases, the anaphoric expression contains additional information. such “specificational” anaphora are known for gender (e.g. how do you think god looks like? – well, i think she is black, where the pronoun she resolves the underspecified gender of the antecedent god to female). proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 227 https://doi.org/10.3765/elm https://www.elm-conference.net/ farkas & de swart (2003), discussing hungarian, claim that pin objects can be picked up by null anaphora. modarresi (2014, 2015) assumes that this is the preferred uptake in persian as well and explains this by null anaphora not having a number feature. but modarresi also allows for singular anaphora (-esh) and plural anaphora (-eshoon), as their semantic restrictions are compatible with the number neutrality of pin antecedents. in these cases the anaphoric expression are specificational, as they carry additional information (that sara bought one book, or that sara bought more than one book). such additional information may be supported by stereotypical interpretations, as in (4) and (5) for the plural and singular interpretation, respectively. hence, the preferred stereotypical interpretation should have an influence on the nature of the anaphoric uptake of pins. according to modarresi (2014, 2015), there is no fundamental difference in the anaphoric potential of yek-marked objects and bare objects in persian; they both introduce a dr that is accessible for future uptake. modarresi does not assume that singular, plural or neutral drs differ in their markedness; if there are differences, we should assume that number-neutral drs are least marked. the second theory we consider is krifka & modarresi (2016). according to it, the event dr is bound by a narrow-scope existential quantifier, the existential closure introduced by diesing (1992), which scopes over the syntactic domain of the vp. objects with a rā scramble out of the vp -hence escape existential closureand have to be interpreted outside of the scope of the existential quantifier (modarresi 2014). yek-marked objects not marked by rā stay within the vp but can scope within or outside of existential closure, a variability that is known for indefinites with determiners in general (cf. fodor & sag 1982). a further assumption that sounds unintuitive initially is that bare nouns are definites, with a singular interpretation. this holds uncontroversially for subjects and for objects marked by rā, which tend to have a definite, number-specific interpretation, cf. (1). but we take bare objects not marked by rā, which remain within the vp, to be singular definites as well. this is possible because we assume that bare nouns in general are dependent definites. the apparent indefinite number-neutral interpretation of bare objects without rā marking arises as a secondary effect due to the place where the bare noun is interpreted, within the scope of the existential quantifier over the event, and as functionally related to the event. one theoretical advantage of his hypothesis is that a uniform interpretation of bare nouns as singular definites as subject, rā-marked objects, and objects that are not rā-marked becomes possible. the narrow-scope indefinite, number-neutral interpretation of bare objects comes about as illustrated in the following examples. the wide-scope interpretation of example (7) with yekmarked object is given in (13) and (14): (13) ⟨ | ⟩ + (7)(a) = (x₁ x₂ )x₁ = sara, ∃⟨e₁| book(x₂), |x₂| = 1, e₁: bought(x₁,x₂)⟩ (14) (13) + (7)(b) = # x₁ x₂ % x₁ = sara, ∃⟨e₁| book(x₂), |x₂| = 1, e₁: bought(x₁,x₂)⟩ ∃⟨e₂| e₂: read(x₁,x₂)⟩ & notice the condition of the form ∃⟨ d | c⟩, where ∃ stands for diesing’s existential closure operator. this is a complex condition, other examples of complex conditions being negation, disjunction, and quantification (cf. kamp & reyle 1993). the condition ∃⟨ d | c⟩ holds with respect to a function g and a model m iff g can be extended to a function g′ that also maps the drs in d to entities in m such that all the conditions in c are satisfied in m. the resulting drs (14) is truth-conditionally equivalent to (9). the interpretation of (10) is given as follows, on the input of an empty drs ⟨ | ⟩. (15) ⟨ | ⟩ + (10)(a) = (x₁ ) x₁ = sara, ∃⟨e₁ x₂| x₂=book-of(e₁)(x₂), |x₂| = 1, e₁: bought(x₁,x₂)⟩ proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 228 https://doi.org/10.3765/elm https://www.elm-conference.net/ here the object is interpreted as a definite dependent on the event e₁, as ‘the unique (single) book of e₁’. as e₁ is introduced within existential closure, the corresponding dr x₂ has to be introduced within existential closure as well. a consequence of this is that x₂ cannot be accessed directly by the following sentence. but kamp & reyle (1993) have proposed, for quite different reasons, that drs in subordinated drss can be reactivated by an operation called abstraction and summation. for the case at hand, this is illustrated in (16). (16) (15) + (10)(b) = 0 x₁ x₂ x₃ 2 x₁ = sara, ∃⟨e₁ x₂| x₂=book-of(e₁)(x₂), |x₂| = 1, e₁: bought(x₁,x₂)⟩ x₃ = σx₂ ⟨e₁ x₂| x₂=book-of(e₁)(x₂), |x₂| = 1, e₁: bought(x₁,x₂)⟩ ∃⟨e₂| e₂: read(x₁,x₃)⟩, |x₃| ≥/=/> 1 6 in the second line, a new dr x₃ is introduced and anchored to the sum (σ) of all x₂ that satisfy the condition expressed in the scope of σ, namely that there is an event e₁ such that x₂ is the unique book of e₁ and e₁ is an event of x₁ buying x₂. this is the sum of all books that sara bought, in the relevant discourse universe. notice that this sum can be one or more than one book. the dr x₃ can be taken up in the second sentence by a dr that is number-neutral (|x₃| ≥ 1) as with null anaphora, or atomic (|x₃|=1) as with the singular anapher -esh, or non-atomic (|x₃|>1) as with the plural anaphor -eshoon. in contrast, in case of a yek-marked indefinite as in (14), only a neutral or singular anaphor is possible. this approach differs from the assumption of number-neutral drs, as it assumes a more complex mechanism of anaphoric update in the case of bare noun objects compared to yek-dps. in general, we should find that anaphoric uptake for bare noun (pin) is less frequent than with yekmarked nouns. this is what modarresi & krifka (to appear a, b) indeed found, in particular in their free sentence completion task (see section 3.3 below). notice that just as for the approach with number-neutral drs in modarresi (2014, 2015), cf. (12), this analysis predicts an influence of stereotypical world knowledge on the use of singular or plural anaphoric devices. if world knowledge suggests that more than one entity was subjected to the event, as in (4), we should easily find the plural anaphor -eshoon next to the null anaphor, but the singular anaphor -esh should not occur. this is in contrast to cases like (5) which suggest that only one entity is involved. but the analysis presented here differs from modarresi (2014, 2015) in one respect. the simplest summation is in case there is only one relevant event in the model, as it then amounts to referring to the single atomic individual that is involved in that event. in this limiting case, the summation operation σ is reduced to the identifying function x₃ = ιx₂⟨e₁ x₂ | …⟩. hence, we should find, in addition to an effect of stereotypical world knowledge, a preference for anaphoric uptake by singular pronouns, instead of null or plural pronouns, as simple summation would yield drs that are anchored to atomic individuals. this is different to modarresi (2014, 2015), who assumes number-neutral drs; in this case, number-neutral null anaphora should be the preferred choice of anaphoric uptake. 3. experimental evidence. in this section we will present three experiments that provide evidence for the hypotheses presented in the previous sections. in order to facilitate the discussion, we list here five hypotheses, where a and a* are related to modarresi (2014, 2015) according to which pins are discourse transparent, and b and b* are related to krifka & modarresi (2016) according to which pins are discourse translucent. 0 is the hypothesis that pins do not introduce drs, hence that they are discourse opaque. proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 229 https://doi.org/10.3765/elm https://www.elm-conference.net/ (17) specific hypotheses: 0: bare noun (bn) objects do not license anaphora. a: bn objects license anaphora to the same degree as yek-marked (yk) objects. b: bn objects license anaphora to a reduced degree compared to yk objects. a*: bn objects introduce number-neutral drs, preferred uptake by null anaphora b*: bn objects introduce drs by existential closure, preferred uptake by singular anaphora experiment 1 involves forced choice of anaphora; experiment 2 forced choice of antecedents; experiment 3 is a free text completion task. hence, all three experiments test language production. 3.1. forced choice of anaphora. in this experiment we tested the anaphoric potential of bare noun (bn) objects and yek-marked (yk) objects in a forced-choice selection of anaphoric expressions, a controlled production experiment. participants were presented with a sentence containing a bn object or yk object in antecedent sentences that were constructed in a way as to have a bias towards a singular or, a plural interpretation, or no particular interpretation (neutral bias). the continuation sentence contained a blank to be filled by null, singular (sg), or plural (pl) anaphora. the reason for testing three types of biases was based on the observation by modarresi (2014) that such biases may affect the choice of pronominal anaphora referring to bn antecedent. we constructed 36 test items including 8 fillers. the test items had 6 conditions: (2 antecedent types and 3 bias types). as a sample item representing all six conditions and the possible three reactions, consider (18). the experimental items were presented in persian script, of course. (18) sara { yek / --} {television / ketāb / havij } kharid. sara idf / bn tv book carrot bought baad tu-ye mashin ▢ gozasht-esh ▢ gozasht ▢ gozasht-eshoon then in-ez car put-it put-ø put-them we list the test items in shortened form for singular bias (19), neutral bias (20) and plural bias (21) in persian together with an idiomatic translation. we assigned the bias category following our own intuition, also asking other native persian speakers. while it is relatively easy to find examples with clear singular or plural bias, the construction of examples with neutral bias are less clear-cut.1 (19) khooneh be ers bord ‘inherited house’, mashin lebasshoui kharid ‘bought washing machine’, motorcyclet kharidam ‘bought motorcycle’, lebase aroos keraye kard ‘rented wedding dress’, baraye doostash sandwich kharid, ‘bought sandwich for a friend’, baraye behtarin danesh-amooz jayezeh sefaresh dad ‘ordered prize for the best student’, docharkhe did ‘saw bicycle’, keike tavallod avard ‘brought birthday cake’, ghayegh ejareh kard ‘rented boat’, gavsandogh dozdid, ‘stole safe’, khooneye jadedemoon piano dareh ‘our new house has piano’, saate tala barandeh shod ‘won golden watch’ 1 the construction of examples could have been objectivized based on a very large corpus. for example, google n-grams shows that the rate of occurrences of the string had a piano vs. had pianos is about 14, of received a gift vs. received gifts is about 1.12, and of corrected a paper vs. corrected papers is about 0.01, indicating a clear single, neutral, and plural bias, respectively. this is based on english; no comparable large corpus for persian is available. proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 230 https://doi.org/10.3765/elm https://www.elm-conference.net/ (20) jabeh avord ‘brought box’, ketāb kharid ‘bought book’, sham roshan-kard ‘blew candle’, abnabat avord ‘brought bonbon’, ketāb khoond ‘read book’, card-postal ferestad ‘sent postcard’, livan kharid ‘bought glass’, hedyeh gereftam ‘received gift’, badkonak kharid ‘bought balloon’, ketāb gharz gereft ‘borrowed book’, too ashpazkhooneh magas did ‘saw mosquito in the kitchen’, az ghafaseh fenjoon biroon avard ‘took cup from the shelf’ (21) adas rikht ‘poured lentil’, zardaloo chidam ‘picked apricot’, baghboonha too-ye-bagh derakht kashtand ‘the gardeners planted trees in the garden’, too dasht shaghayegh daroomadeh ‘in the pasture bloomed tulip’, baraye sakhtane divar ajor sefaresh-dad ‘ordered brick to build wall’, sherkate rahahan barayaerahhaye mokhtalef ghatar kharid ‘the railroad bought train for different roads’, baraye mehmooni sandali sefaresh dadam ‘ordered chairs for garden party’, moallem baraye bachehha ye class jayezeh gereft ‘the teacher got prize for children in the class’, too bazar havij mifrookht ‘sold carrot in bazar’, tamame rooz livan shostam ‘washed glass all day’, tamame rooz varagheh sahih kard ‘corrected paper all day’, tamame hafteh daman dookht ‘sew skirt all week’ there were 357 native persian speakers that voluntarily participated in this experiment using an online survey platform (socsi survey). the stimuli were presented in twelve different lists.2 each list included all six conditions, with an average of four fillers; the items were randomized using latin square design. the results are indicated in figure 1. figure 1: forced choice of sg, null and plural anaphora in sentences with singular, neutral or plural bias and yek-marked antecedent or bn antecedent. y-axis specifies number of items. as participants were forced to select an anaphoric uptake, the experiment cannot distinguish between hypotheses 0 and a/b. but it showed that bias has an effect on the nature of the anaphoric uptake. bn antecedents were taken up most often by sg pronouns under singular bias, and by pl pronouns under plural bias, consistent with hypotheses a* and b*. comparing yk and bn 2 items from the next experiment, antecedent choice, were also included, that is why we randomized in twelve lists as opposed to six lists. it was made sure that no participant saw the same sentence twice. 0 50 100 150 200 250 300 350 400 450 500 sg null pl singular bias 0 50 100 150 200 250 300 350 400 sg null pl neutral bias 0 50 100 150 200 250 300 350 400 sg null pl plural bias yk antecedent bn antecedent proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 231 https://doi.org/10.3765/elm https://www.elm-conference.net/ antecedents, we find that bn antecedents favored null anaphora a bit more under all three bias conditions. this can be interpreted as a slight preference for hypothesis a* over b*. however, we would expect a considerably stronger preference if null anaphora corresponds to bn antecedents in their number feature, as assumed under hypothesis a*. furthermore, under the neutral bias condition, bn antecedents are taken up by sg and null anaphors equally. this would be expected under hypothesis b*, which assumes a structural tendency for simple summation, in contrast to a*. 3.2. forced choice of antecedent. experiment 1 did not show whether the participants favored the use of bns as antecedents of anaphora because the antecedents were fixed conditions in the experimental items. we reversed the design and investigated the choice of antecedents (bn vs. yk objects), when the anaphor in the subsequent sentence is fixed (as nl, sg or pl). like in the previous experiments we had three types of biases. with the exception of the reversal of the design, the stimuli were the same as in experiment 1. this is illustrated in (22). (22) ali ▢ { television / ketāb / havij} / ▢ yek { television/ ketāb / havij} kharid va ali tv book carrot} idf tv book/ carrot} bought and { gozasht-esh / gozasht-ø / gozasht-eshoon } rooy-e-miz. put-it put-ø put-them on-ez-table there were 36 items with 9 conditions (3 anaphor types x 3 bias types), including 8 fillers. the same 357 native persian-speakers as in experiment 1 participated in experiment 2, as the second part of the experiment. the stimuli were presented in twelve different lists. each list included all the nine conditions in randomized order, one of the conditions of each sentence, including an average of four fillers in each list, using a latin square design. after reading the whole sentence, the participants had select the bn or the yk noun as the most appropriate antecedent. results are presented in figure 2. figure 2: forced choice of yek-marked antecedent of bn antecedent with sg, null and pl anaphora and singular, neutral and plural bias. y-axis specifies number of items. we concentrate first on the cases with singular and neutral bias. clearly, yk objects make better antecedents except for pl anaphors, as in this case there would be a semantic clash between 0 10 20 30 40 50 60 70 80 sg null pl singular bias 0 10 20 30 40 50 60 70 80 90 sg null pl neutral bias 0 10 20 30 40 50 60 70 80 90 100 sg null pl plural bias yk antecedent bn antecedent proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 232 https://doi.org/10.3765/elm https://www.elm-conference.net/ the singular-marked antecedent. this supports hypothesis b over a (the hypothesis that bn and yk-marked antecedents have the same frequency for singular/neutral bias and sg/null anaphors is rejected by a chisquare test with p < 0.001). but the experiment also shows that bn objects were actually selected quite often, even for sg and null anaphors. this is evidence against hypothesis 0. bn objects are naturally favored in cases with plural bias, which disfavors yk antecedents for semantic reasons. notice that these semantic reasons outweigh any preference of yk antecedents as anaphors, even in the case of sg pronouns. surprisingly, even with pl anaphors there was a substantial minority of cases with yk antecedents (about 30 vs. 70 for the cases with singular and neutral bias). a possible reason is that the yk objects were interpreted with narrow scope, introducing their dr in the existentially quantified sub-drs, which would be consistent with hypothesis b*. an alternative explanation is that the task of going back in the text, choosing an antecedent, reading the text, choosing the other antecedent, reading the text in this version, and selecting the better option of the two versions was quite complex. it might have led to selecting the yk variant without reading the whole sentene, because this is in general the better antecedent. 3.3. free completion task. in a final experiment, we investigated which anaphoric forms are generated spontaneously in a free completion of a preceding sentence, contrasting bn and ykmarked antecedents. this task does not investigate the reflection of participants about language but rather asks for a natural production task, leaving many more options. in particular, it also leaves open the option of no anaphoric uptake at all. a sample item is (23). (23) leila { yek / --} { television / ketāb / havij } kharid va baad _________________ leila {idf / bn }{ tv book carrot } bought.3sg and then ‘leila bought (a) tv/ book / carrot and then _________________________’ there were 6 experimental conditions (2 antecedent types x 3 bias types). the stimuli consisted of 36 items including 3 fillers, randomized in a latin square design in eight lists. there were altogether 252 participants that took part in an online experiment. participants read sentences in different conditions and were asked to type a suitable continuation. we collected about 330 to 420 data points for each condition, altogether 2256 data points after exclusion of incomplete answers. every sentence was analyzed separately to see if and how the participants referred back to the antecedent object noun. naturally, there was a greater variety in the anaphoric responses. the results in figure 3 visualize nl anaphora, singular anaphoric reference with pronouns or clitics (pro-sing), singular anaphoric reference will full dps (full dp-sing), plural anaphoric reference with pronouns or clitics (pro-plur) and plural anaphoric reference with full dps (full dpplur). associative plurals and reference to kinds were very rare and are not reported here.3 3 to be sure, associative anaphora do occur in persian, as in other languages. for the experimental task of sentence continuation, participants did not employ this device, presumably because it requires the introduction of new drs that are licensed by world knowledge, which requires additional effort. we would also like to remark that an associative anaphora analysis of anaphoric uptake by pronouns is implausible. persian has no grammatical gender, hence pronouns are semantically impoverished compared to languages like german and even english. for this reason, associative anaphora of the type leili got married. he is nice. are impossible in persian. proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 233 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 3: free sentence completion task with yk vs. bn antecedents in sentences with singular, neutral and plural bias. pronominal uptake by null anaphora, sg pronouns, full singular dps, plural pronouns, full plural dps, no reference, kind reference, and associative pro-forms. y-axis specifies percentages. we discuss here the main results from the completion experiment. first, we see in the nr column that bn objects are picked up about half of the time (slightly less so in the singular bias). this is definite evidence against hypothesis 0. uptake is only in a minority of cases by full dps, making it implausible that the uptake is predominantly by associative anaphora. concentrating on the singular and neutral bias situations, we see in the nr column that bn objects are less often picked up by anaphora than yk objects, supporting hypothesis b over a. we also see that bn antecedents do not particularly favor null anaphors, supporting hypothesis b* over a*. uptake of bn antecedents by null and singular anaphora is about equal, arguing against hypothesis a*, 0 10 20 30 40 50 60 null pro-sg full dp-sg pro-pl full dp-pl nr kind associative pro singular bias 0 10 20 30 40 null pro-sg full dp-sg pro-pl full dp-pl nr kind associative pro neutral bias 0 10 20 30 40 50 null pro-sg full dp-sg pro-pl full dp-pl nr kind associative pro plural bias yk antecedent bn antecedent proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 234 https://doi.org/10.3765/elm https://www.elm-conference.net/ which predicts a greater affinity of bn antecedents to null anaphora. this is consistent with hypothesis b*, which states that there is a tendency to prefer simple summation, licensing singular anaphora. one result that is difficult to interpret is why there is more anaphoric uptake in the cases of neutral bias, compared to singular and plural bias. 4. conclusion. in this article we discussed the nature of pseudo noun incorporation (pin), with persian bare noun objects (bn) as example. most theories discuss semantic properties of pseudo incorporation, which includes their unspecificity as to singular or plural interpretation. we argue that the anaphoric potential, the ability to be taken up by anaphoric expressions, is crucial for the proper analysis of pins. we investigated their anaphoric potential in contrast to regular singular indefinites marked with the determiner yek ‘a / one’ in experimental items with three types of biases: bias towards a singular interpretation, towards a plural interpretation, and with neutral bias that neither favors singular nor plural interpretation. despite the widespread assumption that pins do not introduce discourse referents (drs) (for persian bns, cf. modarresi & krifka to appear a), our experimental results have shown that bare nouns are actually quite good antecedents – though slightly less than indefinite antecedents. the focus of the current paper was on the effect of biases towards singular or plural interpretations of pin objects, and the absence of such biases. the experiments provided evidence that can be cautiously interpreted as disfavouring theories that assume that pin objects are semantically specified as number-neutral such as modarresi (2014, 2015), as then we would have found a greater preference for number-neutral null pronouns for pin objects. rather, the experiments favour the proposal by krifka & modarresi (2016), which considers pin objects as dependent singular definites within existential closure over an event variable. the dr of pins is not directly available but can be accessed via a process of abstraction and summation, a phenomenon that is well-known in other cases, as in so-called donkey sentences. the process predicts a general preference for singular drs, for which there is evidence in our experimental data. references aguilar-geuvara, ana & zwarts, joost. 2010. weak definites and reference to kinds. semantics and linguistic theory (salt) 20, 179–196. borik, olga & berit gehrke. 2015. an introduction of the syntax and semantics of pseudo-incorporation. in: borik, olga & berit gehrke (eds.), the syntax and semantics of pseudo-incorporation. leiden: brill, 1-46. brocher, andreas et al. 2020. referent management in discourse: the accessibility of weak definites. cognitive science 2829-2835. chung, sandra & william ladusaw. 2020. noun incorporation. in gutmann, daniel et al. (eds), the william blackwell company to semantics. dayal, veneeta. 2011. hindi pseudo-incorporation. natural language and linguistic theory 29, 123–167. diesing, molly. 1992. indefinites. cambridge, ma: mit press. dobrovie-sorin, carmen & ion giurgea. 2015. weak reference and property denotation: two types of pseudo-incorporated nominals. in: gehrke, berit & olga borik, (eds), the syntax and semantics of pseudo-incorporation. brill, 88-125. espinal, m. teresa & louise mcnally. 2011. bare nominals and incorporating verbs in spanish and catalan. journal of linguistics 47, 87-128. farkas, donka, and henriëtte de swart. 2003. the semantics of incorporation: from argument structure to discourse transparency. stanford: csli publications. proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 235 https://doi.org/10.3765/elm https://www.elm-conference.net/ fodor, janet dean & ivan a. sag. 1982. referential and quantificational indefinites. linguistics and philosophy 5: 355-398. ghomeshi, jila. 2008. markedness and bare nouns in persian. in: simin karimi, vida samiian & donald stilo (eds), aspects of iranian linguistics. cambridge scholar publishing. 85-112. kamp, hans & uwe reyle. 1993. from discourse to logic. introduction to model theoretic semantics of natural language, formal logic, and discourse representation theory. dordrecht: kluwer. kamp, hans, uwe reyle & josef van genabith. 2011. discourse representation theory. in: guenthner, franz & dov m. gabbay, (eds), handbook of philosophical logic. springer, 125-394. krifka, manfred & fereshteh modarresi. 2016. number neutrality and anaphoric uptake of pseudoincorporated nominals in persian (and weak definites in english). salt 26, 874-891. ladusaw, william & sandra chung. 2004. restriction and saturation. cambridge, ma: mit press. law, jess h.-k. & kristen syrett. 2017. experimental evidence for the discourse potential of mandarin. north eastern linguistics socienty (nels) 47. massam, diane. 2001. pseudo noun incorporation in niuean. natural language and linguistic theory 19: 153-197. massam, diane. 2017. incorporation and pseudo-incorporation in syntax. in: aronoff, mark, (ed), oxford research encyclopedia of linguistics. mcnally, louise. 1995. bare plurals in spanish are interpreted as properties. in formal grammar, eds. glyn morrill & richard oehrle (eds.), formal grammar. barcelona: polytechnic university of catalonia, 197-212. mcnally, louise & veerle van geenhoven. 2005. on the property analysis of opaque complements. lingua 885-914. modarresi, fereshteh. 2014. bare nouns in persian: interpretation, grammar, and prosody. doctoral dissertation. humboldt universität zu berlin. modarresi, fereshteh. 2015. discourse properties of bare noun objects. in olga borik & berit gehrke (eds.), the syntax and semantics of pseudo-incorporation. leiden: brill, 189-221. modarresi, fereshteh & manfred krifka,. 2020. anaphoric potential of pseudo incorporated nominals in comparison with compounds and implicit objects. linguistic evidence 2020. modarresi, fereshteh & manfred krifka. to appear a. anaphoric potential of pseudo-incorporated bare objects in persian. in simin karimi (ed), north american conference on iranian languages (nacil2). modarresi, fereshteh & manfred krifka. to appear b. pseudo-incorporation and anaphoricity: evidence from persian. glossa. schwarz, florian 2014. how weak and how definite are weak indefinites? in: aguilar-guevara et al. (eds), weak referentiality. amsterdam: john benjamins. van geenhoven, veerle. 1998. semantic incorporation and indefinite description. standford: csli publications. yanovich, igor. 2008. incorporated nominals as antecedents for anaphora, or how to save the thematic arguments theory. university of pennsylvania working papers in linguistics 14, 367-379. proceedings of elm 1: 224-236, 2021 fereshteh modarresi and manfred krifka: pseudo-incorporated antecedents and anaphora in persian: the influence of stereotypical knowledge. 236 https://doi.org/10.3765/elm https://www.elm-conference.net/ clause types and speech acts in speech to children anissa zaitsu, jad wehbe, valentine hacquard & jeff lidz* abstract. the question of how and when children learn to associate clause type with its canonical function, or speech act, is currently unknown. it is widely observed that declaratives tend to result in assertions, interrogatives in questions, and imperatives in requests. although these canonical links between clause type and speech act are principled, they are known to be defeasible. in this corpus study, we investigate how parents talk to their children in the first years of life, and ask how their input might support this mapping, and to what extent it might pose difficulties. we find that the expected link between clause type and speech act is robust in the input, particularly between declaratives and assertions, both of which also occur most frequently. in addition, the non-canonical mappings that do occur are characterized formally, e.g., non-interrogative questions nearly always exhibit rising prosody, and non-imperative requests often contain a modal. keywords. corpus-study; pragmatics; syntax; acquisition; clause types; speech acts 1. introduction. clause types are said to exist in all languages, and are determined either morphologically, syntactically, or in combination with one another, forming the basic categories: declarative, interrogative, and imperative (sadock & zwicky 1985). each clause type has what some have called a ‘default’ function in context (roberts 2018): declaratives typically express assertions, interrogatives typically pose questions, and imperatives typically issue a directive, or what we neutrally refer to as a request. while principled links between clause types and speech acts exist, the correlation is defeasible; for example, a declarative might make a request, or pose a question. given this, one might ask when and how children come to identify the clause types of their language, and their default functions. the literature does not provide a clear answer to this. we know that by 12 months children can distinguish formal properties of the clause types declarative and interrogative (e.g., geffen & mintz 2015) and by age 3, have associated clause type with its canonical speech act (e.g., rakoczy & tomasello 2009). but when exactly does this knowledge come on line, and how? previous work suggests that distinguishing clause types may be crucial for the acquisition of basic syntactic properties of one’s language, argument structure and word meanings before their second birthday (pinker 1984, pinker 1989, gleitman 1990, gleitman et al. 2005, perkins 2019). knowledge of declarative clauses, in particular, would be useful early on, as they are ‘basic’ in a sense, in that they might establish a baseline with respect to word order and argument structure in a language, which is certainly true of english. figuring out what the relevant clause types are, however, is not trivial. languages typically employ a wide range of clausal constructions—actives, passives, imperatives, polar and wh-interrogatives, clefts, pseudoclefts, relative clauses, and so on. nonetheless, many of these distinctions are collapsed from the perspective of clause typing. for *we would like to thank dan goodhue, laurel perkins, alexander williams, and yu’an yang for helpful discussion and feedback. authors: anissa zaitsu, stanford university (azaitsu@stanford.edu), jad wehbe, mit (jadwehbe@mit.edu), valentine hacquard, university of maryland (hacquard@umd.edu) & jeff lidz, university of maryland (jlidz@umd.edu). proceedings of elm 1: 284-297, 2021 c©2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz published by the lsa with permission of the author(s) under a cc by license. 284 https://doi.org/10.3765/elm https://www.elm-conference.net/ instance, verbs that select for declarative complements do not distinguish between actives, passives and clefts. children, then, need to cluster the right morphosyntactic properties that are relevant for clause types and associate them with the right pragmatic function. to start addressing how children might make this association, we investigate how parents speak to their children through a corpus study that examines both the range of clause types and speech acts parents use, and ask whether the two align in the ways we expect. we focus on ages 1 to 3, since this is the period during which children likely associate clause type and speech act, to see what components of the input support this mapping, and identify where there might be bottlenecks. we ask, in particular, whether parents’ speech changes over time, as children’s language develops, and they become increasingly able to advance and recognize conversational goals. for example, parents might be less likely to ask questions when children can’t respond, which could affect when children link interrogatives to questions. our results show that the mapping between clause type and speech act is relatively stable in the child’s input; each clause type is mostly used for its default function. we find that overall, declaratives are the most frequent clause type, and that assertions are virtually always made with a declarative. together, this suggests that children could identify declaratives as basic early on, which would aid them in the ways alluded to above. questions and requests, compared to assertions, are more varied in terms of the forms that express them. however, in such cases, we find stable indicators of a mismatch. for example, rising prosody tends to mark non-interrogative questions, and modals tend to mark non-imperative requests. finally, the rate at which each clause type and speech act occur across the ages also remains relatively stable, except for wh-questions, which surprisingly, occur more in the first year of life than in both the second and third. the input then supports a view on which children use the link between clause types and speech acts to learn the form and function of the clause types. 2. background. one of the difficulties in formal work on clause types and speech acts is providing the right semantics, usually referred to as the mood, for each clause type such that clause type doesn’t fully determine speech act (roberts 1996, portner 2004, farkas & bruce 2010, starr 2014, krifka 2017, murray & starr 2020, a.o.), but that it still lets us understand why we find the principled links that we do. for example, while a declarative typically makes an assertion, it might also make a request: “you have to leave,” and famously, interrogatives, which typically pose a question, might also function as a request, “can you pass the salt?” (searle 1975). a prominent view of speech acts holds that the distinctions among each speech act lies in what they propose to their interlocutors. on roberts’ (2018) view, an assertion involves proposing adding the content of p (some proposition) to the common ground (stalnaker 1978). a request proposes that interlocutors adopt intentions to make p true. a question proposes that the interlocutors collectively commit to collaborative inquiry with respect to p. the kinds of proposals each clause type can make are constrained by the semantic content of p, and the mood operator characteristic of each clause type. the declarative mood denotes a proposition, which naturally lends itself to assertions, which assert p. interrogative clauses denote a set of propositions, which explains why they invite collaborative inquiry over a set of alternatives. the semantic type for imperatives has been more widely debated, but one view, which will suffice for our purposes is that they denote a property indexed to the addressee and targets the priorities of that addressee (portner proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 285 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2004, 2007, roberts 2004, starr 2020). it has been established that by age 3, children have properly linked declarative clauses with assertions, imperatives with requests, and that they properly answer interrogatives, suggesting an understanding of them (rakoczy & tomasello 2009). still, it is unclear how and when they make this link between clause type and speech act. we know that very early on, infants can track the state of knowledge of the people around them, as well as their goals and intentions, including their referential intentions (woodward 1998, onishi & baillargeon 2005, kovács et al. 2014, martin et al. 2012). we also know that young infants show remarkable abilities to track formal properties of the utterances they hear (saffran et al. 1996, gómez & gerken 2000, mintz 2002). through distributional analysis, the child can identify clusters of syntactic features that often co-occur and are good indications of clause type. geffen & mintz (2015, 2017) argue that formally, children distinguish declaratives from interrogatives from around 12 months. but children need to further figure out which clause type is which, that is, its primary function. this is where speech act information could be crucial. we want to know what information children have available to them to make this link, which is in place by 3 and what, if any, difficulties might arise as a function of their input. for example, we know that the link between clause type and speech act is not absolute, but we don’t yet know to what extent non-canonical mappings actually occur in natural speech to children. were they to be very frequent, this could pose a difficulty for the child trying to establish each clause type’s basic function. we also want to know whether utterances which do not conform to the default clause type speech act mapping have regularities which the child might be able to track and recognize as exceptions to the rule. 3. procedure. we examined 15,243 parent utterances from the providence portion of childes (demuth et al. 2006) from ages 1;00 to 3;07. this consisted of 5 parent-child pairs who were recorded in both audio and visual formats during unstructured play-time regularly over a period of three years. for each pair, 1,000 parent utterances were annotated per age group: 1;00-1;11 (“1-year olds”), 2;00-2;11 (“2-year olds”), 3;00-3;11 (“3-year olds”), which resulted in roughly 3,000 utterances per parent-child pair. each utterance was marked for clause type: declarative, interrogative, and imperative. (1) a. johnny picked the flowers. declarative b. did johnny pick the flowers? interrogative c. what did johnny pick? interrogative d. pick the flowers. imperative declaratives are marked in english by the presence of a subject and agreement on the verb, interrogatives are marked by subject-auxiliary-inversion (sai) or by a left-edge wh-phrase, and imperatives are marked by the absence of a subject and no tense marking on the verb. in addition, the marginal clause type exclamative was used to mark wh-exclamatives (reminiscent of exclamatives described in rett 2011): (2) what a super idea! exclamative proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 286 https://doi.org/10.3765/elm https://www.elm-conference.net/ in naturalistic data, there are also many utterances which do not meet any of the criteria just described. we used the tag frag as a catch-all for sub-sentential utterances: (3) a. good boy. frag b. nice of you. frag c. alright like that. frag in addition to clause type, we also marked utterances for speech act. speech acts were coded according to what the annotator intuited the parent to be doing with the utterance. we view speech acts as proposing to update the context in a particular way. that is, utterances which attempt to give the child information were marked as assertion; utterances that offer the child some sort of direction were marked as request; utterances which invite collaborative inquiry were marked as question; and utterances which express some degree of surprise for p were marked as exclamation. each utterance was coded by one of the two annotators (the first authors on this paper). 10% of the corpus coded by both annotators to check for reliability. we checked for inter-annotator agreement on speech act, which was high (cohen’s kappa= .92). the rest of the corpus was coded by one of the two annotators. because of the audio/visual format of the corpus, it was easy to detect intonation, and to see how the child reacted to the parents’ utterances, providing a way to reasonably determine the intended speech act. below are examples of what we would call a match; that is, the clause type used results in the expected speech act: (4) a. is this turned on? interrogative-question b. go get the other one. imperative-request c. that’s the photograph book we made. declarative-assertion d. what a good fish face! exclamative-exclamation certain frag utterances were also marked as having a speech act if the annotator was able to discern a full vp; that is, a predicate with its internal argument or a bare v if intransitive. (5) a. saw a lot of them earlier. frag-assertion b. see the shovel? frag-question c. may have to do it to other way. frag-request (5-a) is missing a subject, but the past tense marking clearly rules it out as an imperative clause. in (5-b), a subject is missing once again, and while there’s no tense marking on the verb, the utterance is marked by rising intonation. the verb in (5-b) is also the perception verb see, which does not have an agent and cannot be used as an imperative. in (5-c), a subject is missing, but there are two modals above the vp, which means it cannot be an imperative clause, and provides enough information to see that (5-c) expresses a request. we also find variety of mismatches involving full clauses, a subset of which are presented below. a mismatch is defined as a clause that does not result in its default speech act. (6) a. you have to go forwards. declarative-request b. can you move your foot please? interrogative-request c. oh my gosh look at all those babies! imperative-exclamation proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 287 https://doi.org/10.3765/elm https://www.elm-conference.net/ the imperative-exclamation seen in (6-c) wasn’t very common, but it’s clear in context that they do not direct their addressee, and instead meet the criteria for exclamation. in the context that (6-c) occurs, the child is showing the parent several baby dolls, which does not support a directive use of the imperative form. in addition to clause type and speech act information, annotators also coded for various grammatical aspects of the utterance. various features of the subject were marked, as were modals, embedding verbs, and negation. this was done in order to trace in regularities in mismatches, to be discussed in the following section. lastly, utterances were also marked for intonation – this helped to identify questions across clause type. the form of polar questions and the function of wh-elements proved to be fairly diverse in the input, which required further elaboration of the annotation schema. polar questions are those which ask whether p or its negation holds. although rising declaratives (rd) have a similar effect in terms of introducing alternatives p and its negation, they are said to introduce an additional bias for one of the alternatives (gunlogson 2008, malamud & stephenson 2015, farkas & roelofsen 2017). whether they are fundamentally declarative in mood or interrogative has recently been taken up (farkas & roelofsen 2017, malamud & stephenson 2015, jeong 2018). we do not take a theoretical stance here, but we did mark them with declarative clause type and additionally marked them for rising intonation – they were also given a speech act, which usually (but not always) turned out to be question, and thus were counted as a mismatch. however, under certain approaches (e.g., farkas & roelofsen 2017), characterizing a rd as a mismatch wouldn’t be wholly accurate, since rising declaratives are viewed as having interrogative mood. in addition, there were several truncated clauses that had a polar question type meaning and the associated rising intonation, which we marked as left edge ellipsis (lee). in some cases, only an auxiliary was elided, which left the subject and the vp. in others, both a subject and an auxiliary were missing, leaving the vp. some were ambiguous between lee and rd, which we called rd-poss. this occurred when there was a 2nd person subject followed by an untensed verb. in english, agreement with a 2nd person pronoun is null, so it’s not possible to tell whether an auxiliary was elided at the left edge or not. (7) a. you’re going to open the door? rd-def b. you want your ball? rd-poss c. you kissing that baby too? lee d. wanna go on your swing? lee e. do you want your ball? pq in (7-a), the auxiliary expressed in its contracted form clearly marks the clause as a rising declarative (rd-def), since we can see that no sai has occurred. on the other hand, the string in (7-b) corresponds to possible declarative clause in english, but the absence of an auxiliary raises an ambiguity. (7-b) could underlying be (7-e), with auxiliary do elided at the left edge, hence the tag rd-poss for possible rd. that such a possibility is available in the language is evidenced by the existence of examples such as (7-c). the -ing marking on verbs is only licensed by a particular kind of auxiliary be in english, i.e., you kissing that baby too is not a possible declarative in english. in that case, we can be sure that some element which is doing the licensing for that form is not pronounced. the same process could be at work in (7-b). we marked both rd-def and rd-poss proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 288 https://doi.org/10.3765/elm https://www.elm-conference.net/ as having declarative clause type, and we marked both kinds of lee as frag. wh-questions (wh-q), on the other hand, were typically expressed as full clauses. there were a minority of fragment utterances that contained a wh-element, however. there were what some might call ‘root sluices’ – that is, a wh-phrase related to some correlate in an antecedent provided in context. there was also a conventionalized ‘what?’ where the parent generally meant, what did you say?, or a conventionalized ‘which one?’ meaning which one are you talking about? also surfaced were what-about-dp and how-about-xp utterances that seem to be conventionalized to direct the child’s attention to some object or to make a suggestion. although a full vp was not always present in these cases, the contribution of the conventionalized aspect of these and the wh-element used made it possible to discern various speech acts. for that reason, fragment wh-questions were also given a speech act. (8) a. how about putting them on this side now? frag-request b. what about elephants? frag-question c. what, honey? (=what did you say?) frag-question d. which one? (=which one are you talking about?) frag-question very little work currently exists on what-about and how-about type questions, so it’s hard to say anything about how to interpret these data. but they seem to have a rather flexible discourse use, usually functioning as a suggestion, a question, and sometimes, even, an answer to a question. 4. corpus data. 4.1. input by age. the data collected allow us to look at various syntactic, semantic, and pragmatic features of the child’s input. we can start by establishing a baseline for the clause types children hear across each age group, allowing us to track whether the input changes. in table 1 below, we divide declarative by fall and rise to indicate intonation, and interrogative by wh and polar. declarative interrogative imperative total fall rise wh polar 1 year olds 1281 (45.2%) 134 (4.7%) 605 (21.3%) 304 (10.7%) 512 (18.1%) 2836 (100%) 2 year olds 1806 (55.7%) 199 (6.1%) 362 (11.2%) 373 (11.5%) 502 (15.5%) 3242 (100%) 3 year olds 1821 (57.5%) 151 (4.8%) 339 (10.7%) 306 (9.7%) 550 (17.4%) 3167 (100%) total 4908 (53%) 484 (5.2%) 1306 (14.1%) 983 (10.6%) 1564 (16.9%) 9245 (100%) table 1: clause type by age the table above, in part, confirms the intuition that declarative clauses are ‘basic.’ across all ages, declarative clauses make up the majority of full clauses. wh-questions are the next most common clause type only for 1-year-olds, making up 21.3% of the full clauses they hear in this sample. that number decreases by 10% by the second year of life. conversely, the rate of falling declaratives increases by 10% by the second year of life. otherwise, the numbers across the other clause types stay relatively stable across all ages. we can also view how speech acts changed across the ages, which largely tracks the changes proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 289 https://doi.org/10.3765/elm https://www.elm-conference.net/ we see in the clause types, foreshadowing that the two are dependent, as expected. below, we show table 2, which takes into account all speech acts, even those with the frag clause type. we also include distinctions among the different kinds of questions: wh, polar, and those which were either of the form rd or lee, which might function differently in their associated context of utterance from basic polar questions made with interrogative syntax. assertion question request total wh polar rd/lee 1 year olds 1255 (40.8%) 650 (21.2%) 230 (7.5%) 236 (7.7%) 701 (22.8%) 3072 (100%) 2 year olds 1736 (50.7%) 410 (12%) 321 (9.4%) 251 (7.3%) 707 (20.6%) 3425 (100%) 3 year olds 1770 (53%) 395 (11.8%) 265 (7.9%) 214 (6.4%) 694 (20.8%) 3338 (100%) table 2: speech act by age we find, also, that there is general stability in terms of speech acts across the ages. assertions are always the most frequent in the child’s input, which seems to track with declarative across the ages. assertion increase by %10 in speech to 2 years old versus speech to 1 year olds, while wh-questions decrease by 10% during this period. taking all the various question types together, they are the next most common speech act, making up about 36% of the input at age 1, 29% at age 2, and 26% at age 3. thus, it does not seem to be true that parents use fewer questions in the earlier years of life, even though children’s ability to accurately understand and respond may not yet be in place. we see, instead, that the opposite is true. parents use more questions in the earlier years of life, and use more assertions as time goes on. 4.2. clause type and speech act mapping. we can now turn to the mismatches between clause type and speech act, which can be viewed in two ways. first, we can assess the rate at which each speech act was performed by each clause type. that is, how often an assertion was made with a declarative, or a question with an interrogative, and so on. we expect to see a tight link between clause type and speech act, given the view that each clause type has a particular semantic make up that gives rise to tendencies among which speech acts they perform. we see this borne out in table 3, below. assertion question request exclamation declarative 4643 (97.4%) 458 (15.3%) 289 (15.3%) 2 (5%) interrogative 18 (0.4%) 2074 (69.2%) 195 (9%) 2 (5%) imperative 0 0 1543 (73.8%) 21 (52.5%) exclamative 0 0 0 13 (32.5%) frag 100 (2%) 440 (15.4%) 75 (2.9%) 2 (5%) total 4761 (100%) 2972 (100%) 2102 (100%) 40 (100%) table 3: percentage of each clause type given the speech act the most stable link from this perspective is that between speech act assertion and declarative. assertions are not readily performed by any of the other clause types. this data point, once proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 290 https://doi.org/10.3765/elm https://www.elm-conference.net/ again, seems to confirm that declarative clauses are ‘basic,’ and should be easily acquirable from the input, particularly if one holds the view that children use this information to learn the clause types of their language. the very few utterances marked as interrogative-assertion either involved a question with high preposed negation, indicating speaker bias for p, or involved an embedded clause, which seemed to be asserted in a clause that was polar in the root clause. (9) a. isn’t that nice? b. did you know that some birds can learn to say words? in these cases, the annotator judged the parent to be proffering the p, that’s nice in (9-a), and to be proffering the p of the embedded clause in (9-b), which meets the criteria for assertion. it should be noted that not all instances of biased polar questions, nor instances of embedded propositions under a root polar interrogative were judged as assertion – these are embedded within particular contexts that gave rise to such judgements. we see, also, of course, that the other core speech acts, question, and request are associated significantly with their expected clause type, though with a greater degree of variability than assertion. we will discuss the properties that characterize these mismatches individually in the following sections. but we should note here before moving on that declarative make up 15% of all mismatches involving question, and another 15% of all mismatches involving a request, foreshadowing that while an assertion is the least variable speech act, declarative might be the most variable clause type. to address this, we can also look at the percentages from the perspective of the clause type. in other words, how often was each clause type used to perform the various speech acts? the data presented below in table 4 will reveal an interesting asymmetry. on one hand, an assertion is rarely made with any other clause type but a declarative, and on the other, declarative turn out to be the most diverse in terms of the speech acts they express. assertion question request exclamation total declarative 4643 (86.1%) 458 (8.5%) 289 (5.4%) 2 (.03%) 5392 (100%) interrogative 18 (0.9%) 2074 (90.5%) 195 (8.3%) 2 (.9%) 2289 (100%) imperative 0 0 1543 (98.7%) 21 (1.3%) 1564 (100%) exclamative 0 0 0 13 (100%) 13 (100%) frag 100 (16.5%) 440 (72.8%) 75 (10.2%) 2 (0.3%) 617 (100%) table 4: percentage that a clause type resulted in each speech act the link between declarative and assertion is still significant from this perspective, but flips to being the most variable cell of the table. the next most variable clause type is interrogative, which result in a request in 8.3% of utterances. the most stringent clause type is an imperative, which nearly always results in a request – the few exceptions being imperativeexclamation instances discussed earlier. proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 291 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4.3. requests. in this section, we take a closer look at what characterizes a request that was not made with an imperative. requests were made with an imperative about 74% of the time. the other 26% were split between declarative clauses, which amounted to 15% of the requests, and another 9% were made with an interrogative. one pattern we identify in table 5 below is that most non-imperative requests contain a modal. declarative total interrogative total mod no mod mod no mod assertion 838 (18%) 3805 (82%) 4643 (100%) 2 (11%) 16 (89%) 18 (100%) question 23 (5%) 435 (95%) 458 (100%) 233 (11%) 1841 (89%) 2074 (100%) request 169 (58%) 120 (42%) 289 (100%) 143 (73%) 52 (27%) 195 (100%) table 5: modals across clause type and speech act utterances marked as declarative-assertion, which constitute a match, only contain a modal 18% of the time. a declarative-question also very rarely contains modals. however, most declarative-request contain a modal, at 58%. this suggests that modals are a good predictor for this type of mismatch, and if tracked by the child, might provide a cue that the clause they encountered has a special function, i.e., is not an assertion. the other 42% of declarative-request without a modal are usually made with a 2nd person pronoun (82%), and with no tense marking on the verb. (10) a. you just give it a little dip. b. you take it with you. c. you line them up. annotators did not specifically code for tense marking, but it’s possible that this, in combination with 2nd person pronouns is also a good predictor of declarative-request. however, one would need to see how often such a frame occurred in declarative-assertion, and so on. table 5 also shows the distribution of modals in interrogative clauses, where modals occur in 73% of those marked interrogative-request. the other two potential mappings skew the other way, in mostly not containing modals. so while interrogative-request make up a relatively small percentage of the mismatches involving request, the majority of them have a modal that could potentially be tracked – and that same feature would be useful in tracking mismatches involving a declarative as well. the other 27% of those marked interrogative-request that do not contain modals also display a particular property. nearly all of them (81%) are a special kind of why-interrogative that involve negation, which together, have a conventionalized meaning: (11) a. why don’t you use the other bat? b. so why don’t you give it to me? c. why don’t you stay right there? negation is relatively rare in the input overall, and most wh-questions that have negation are of the above kind, which means that these should stand out to the child. proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 292 https://doi.org/10.3765/elm https://www.elm-conference.net/ what we have here is good evidence that while mismatches do occur between clause type and speech act, good indicators exist that the clause encountered might have a special function; i.e., they tend to deviate from the patterns displayed in a match. modality, in particular, seems to be especially important in tracking mismatches in both kinds of non-imperative requests. it is also true that in formal work, imperative clauses have been identified with modals more broadly, and have even been said to part of the semantic content of the imperative clause type (kaufmann 2012). from that standpoint, it is then, not surprising that overt modals occur in clauses that ultimately result in a directive force. 4.4. questions. in this section, we take a closer look at the most variable speech act, question. there were several non-interrogative clauses that resulted in a question. 15% were made with declarative clauses, and the other 15% were made with frag utterances. utterances marked declarative-question nearly always had rising intonation – 98% to be exact. the remaining 2% of utterances marked declarative-question were either instances of clauses with an embedded question that were judged as being placed under discussion, or declarative utterances with tags that were judged to be inquisitive in nature. for example: (12) a. i wonder who these are. b. i don’t think today, will we? embedded questions and utterances with tags were not always marked as question. it was in these few contexts that annotators judged such structures to be used as a question. the majority of those marked as frag-question had rising intonation, and were tagged for left edge ellipsis (lee), accounting for 58% of such utterances. the other 42% were instances with a wh-element that wasn’t attached to a full clause, though had recoverable meaning in context. polar questions, in particular, seem to be most variable in terms of the kinds of clauses that express them. in table 6 below, we will see that less than half of all clauses with polar question meaning, i.e., asking whether p or its negation holds, were made with polar interrogative syntax (e.g., displayed sai). q-type example count(%) pq are you thirsty? 808 (48.5%) lee a. wanna go on your swing? b. you kissing that penguin? 256 (15.3%) rd-poss you put bob in the pilot? 242 (14.5%) rd-def you’re gonna hit it or throw it? 202 (12.1%) tag a. we’re not really going to the zoo, are we? b. a dumptruck’s on the tracks, huh? c. you knew that one, right? 158 (9.4%) total 1666 (100%) table 6: polar questions in the child’s input of those utterances that were marked as having a polar question meaning, only 48.5% of them proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 293 https://doi.org/10.3765/elm https://www.elm-conference.net/ were in the full polar interrogative form seen in the first row of table 6.1 if we collapsed the two rd types, rd-poss and rd-def, rising declaratives make up 26.6% of polar questions. instances of lee make up another 15%. most instances of lee were of the vp type seen in (a) of the second row – 173 (67%); the rest were unambiguously missing an auxiliary at the left edge. together, rd forms and lee forms can be unified as having rising intonation, which in turn, unifies them with basic polar interrogatives. despite the variability in morphosyntactic form, children should have a way of understanding these utterances as a unified class. however, some languages routinely use rising intonation to indicate polar question meaning, e.g., french. at some point english-learning children need to learn that they are english, and not french. this may not be a trivial task, given that rising declaratives are fairly robust in the input. the function of a rising declarative is generally thought to be somewhat distinct from both a polar question and an assertion, which means such nuances in force will need to be acquired for english-speaking children. goodhue et al. (2021) show that by age 3, children seem to treat rising declaratives differently from both polar questions and falling declaratives. 5. discussion. our corpus study has shown that the mapping between clause type and speech act is fairly stable in speech to children. each clause type is mostly used to express its default speech act; for the most part, declaratives result in assertions, interrogatives in questions, and imperatives in requests. the same is true from the perspective of the speech act; that is, assertions are nearly always made with a declarative, questions are mostly made with interrogatives, and requests are mostly made with imperatives. we do find variability introduced from this perspective; that is, non-interrogative syntax makes up 31% of questions, and non-imperative syntax makes up 27% of requests. we are able to identify, however, that such uses are characterized by certain formal properties. non-interrogative syntax that results in questions either have rising prosody or a conventionalized use of a bare wh-element. non-imperative requests usually have a modal, or a 2nd person pronoun that is subject of a non-past clause. across the age groups, 1 year olds, 2 year olds, and 3 year olds, the input stays relatively stable from the perspective of clause type. one caveat is that between 12 months and 24 months (1 year olds), children hear wh-questions at a greater rate than in any other age range explored here. this is surprising given that their capacity to respond is fairly limited during this time. in that same age range, declarative clauses are used the least relative to the other two age groups. even so, falling declaratives are the most numerous clause type across all ages. with this much established, we can now address how children might come to associate clause type with its canonical function. one hypothesis consistent with our results is that children use the mutually constraining information from morphosyntactic properties of the clause type and sensitivity to speaker intention, e.g., the speech act, to learn the mapping between clause type and speech act. we know independently from the acquisition literature that children track the formal properties of utterances they hear (saffran et al. 1996, gómez & gerken 2000, mintz 2002). children are also sensitive to the goals and intentions of their interlocutors (woodward 1998, 1to be clear, table 6 is intended to represent utterances with polar question meaning, i.e., were marked as question. that means it does not include the forms interrogative, lee, or rd that were marked as request or assertion. however, because tags were so infrequent overall, we included all such instances in the count above, no matter the speech act they were judged to have. proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 294 https://doi.org/10.3765/elm https://www.elm-conference.net/ onishi & baillargeon 2005, kovács et al. 2014, martin et al. 2012), which could be used to help them track speech act. expectations that a language will generally linguistically distinguish the three main speech acts discussed here would likely be useful, and given the universality of clause types and their basic contextual effect, this seems like a plausible prior expectation. children might further expect that assertions, or utterances that give them information are basic, and that the clause canonically associated with them is, in turn, the most basic clause type, from which they could draw generalizations about word order, argument structure, and recognize various displacements in the language. our results suggest that polar questions might be the most difficult to acquire. less than half of polar questions are made with full interrogative syntax. it is relevant that both kinds of rising declaratives (unambiguous and ambiguous) make up about 26% of the polar questions they hear. in some languages, intonation plus declarative syntax is one way, and sometimes, the preferred way to form a polar question, e.g., french, hindi, brazillian portuguese. children might mistakenly think that rising declaratives are another way to form a basic polar question. in english, however, rising declaratives are said to have a special meaning that distinguishes them in some way from polar questions expressed with interrogative syntax. for instance, one one view, rds are associated with an additional pragmatic effect that combines with the basic contextual effect of an interrogative (i.e., introducing alternatives), creating what often feels like a bias for one of the alternatives introduced (farkas & roelofsen 2017). goodhue et al. (2021) show that 3 year olds treat polar interrogatives, falling declaratives, and rising declaratives differently. whether children can make this three way distinction earlier is not yet known, but given the prevalence of rds in the input and that they may be used where a polar interrogative would also be appropriate, we might expect that at some point along the acquisition trajectory, english-learning children treat rising declaratives and polar interrogatives the same. it’s possible also, however, that the frequency with which polar interrogatives (983) occur as opposed to rising declaratives (484) is enough for the child to determine that rising declaratives are a special kind of polar interrogative, which forms the basic case. 6. conclusion. in this corpus study, we have examined the information that children have available to them in establishing a mapping between clause type and speech act. before children have fully mastered the syntax of their language, sorting clause types would likely be useful, particularly knowledge of declarative clauses as basic. but how do they do this with a limited portion of their grammar in place? our results present a possible way forward. we find a particularly strong link between assertion and declarative clauses. were children equipped with the ability to track speech acts, with prior knowledge that assertions would correspond to the basic clause type in their language, their input would support establishing the mapping early on. among the other two core mappings, the results are slightly more varied, but our results indicate that non-canonical instances are also principled. this could be useful in two ways: i) children might be able to set aside mismatches because they contain markers that indicate they have an effect different from that typically associated with that clause type, and ii) children might be able to learn the special pragmatics associated with these uses, e.g., rising declaratives are a special kind of biased polar question. the study has also raised some new questions, particularly regarding how english-learning children make sense of the varied forms that polar questions take in the language. proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 295 https://doi.org/10.3765/elm https://www.elm-conference.net/ references demuth, katherine, jennifer culbertson & jennifer alter. 2006. word-minimality, epenthesis and coda licensing in the early acquisition of english. language and speech 49(2). 137–173. https://doi.org/10.1177/00238309060490020201. farkas, donka f & kim b bruce. 2010. on reacting to assertions and polar questions. journal of semantics 27(1). 81–118. https://doi.org/10.1093/jos/ffp010. farkas, donka f & floris roelofsen. 2017. division of labor in the interpretation of declaratives and interrogatives. journal of semantics 34(2). 237–289. https://doi.org/10.1093/jos/ffw012. geffen, susan & toben h mintz. 2015. can you believe it? 12-month-olds use word order to distinguish between declaratives and polar interrogatives. language learning and development 11(3). 270–284. https://doi.org/10.1080/15475441.2014.951595. geffen, susan & toben h mintz. 2017. prosodic differences between declaratives and interrogatives in infant-directed speech. journal of child language 44(4). 968–994. https://doi.org/10.1017/s0305000916000349. gleitman, lila. 1990. the structural sources of verb meanings. language acquisition 1(1). 3–55. https://doi.org/10.1207/s15327817la0101 2. gleitman, lila r, kimberly cassidy, rebecca nappa, anna papafragou & john c trueswell. 2005. hard words. language learning and development 1(1). 23–64. https://doi.org/10.1207/s15473341lld0101 4. gómez, rebecca l & louann gerken. 2000. infant artificial language learning and language acquisition. trends in cognitive sciences 4(5). 178–186. https://doi.org/10.1016/s13646613(00)01467-4. goodhue, daniel, jad wehbe, valentine hacquard & jeffrey lidz. 2021. the effect of intonation on the illocutionary force of declaratives in child comprehension. in proceedings of sinn und bedeutung, vol. 25, 1–18. gunlogson, christine. 2008. a question of commitment. belgian journal of linguistics 22(1). 101–136. https://doi.org/10.1075/bjl.22.06gun. jeong, sunwoo. 2018. intonation and sentence type conventions: two types of rising declaratives. journal of semantics 35(2). 305–356. https://doi.org/10.1093/semant/ffy001. kaufmann, magdalena. 2012. interpreting imperatives. springer, studies in linguistics and philosophy 88. https://doi.org/10.1007/978-94-007-2269-9. kovács, ágnes melinda, tibor tauzin, ernő téglás, györgy gergely & gergely csibra. 2014. pointing as epistemic request: 12-month-olds point to receive new information. infancy 19(6). 543–557. https://doi.org/10.1111/infa.12060. krifka, manfred. 2017. negated polarity questions as denegations of assertions. in contrastiveness in information structure, alternatives and scalar implicatures, 359–398. springer. https://doi.org/10.1007/978-3-319-10106-4 18. malamud, sophia a & tamina stephenson. 2015. three ways to avoid commitments: declarative force modifiers in the conversational scoreboard. journal of semantics 32(2). 275–311. https://doi.org/10.1093/jos/ffu002. martin, alia, kristine h onishi & athena vouloumanos. 2012. understanding the abstract role of speech in communication at 12 months. cognition 123(1). 50–60. proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 296 https://doi.org/10.3765/elm https://www.elm-conference.net/ https://doi.org/10.1016/j.cognition.2011.12.003. mintz, toben h. 2002. category induction from distributional cues in an artificial language. memory & cognition 30(5). 678–686. https://doi.org/10.3758/bf03196424. murray, sarah e & william b starr. 2020. the structure of communicative acts. linguistics and philosophy 1–50. https://doi.org/10.1007/s10988-019-09289-0. onishi, kristine h & renée baillargeon. 2005. do 15-month-old infants understand false beliefs? science 308(5719). 255–258. https://doi.org/10.1126/science.1107621. perkins, laurel. 2019. how grammars grow: argument structure and the acquisition of non-basic syntax: university of maryland, college park dissertation. pinker, steven. 1984. language learnability and language development. harvard university press. pinker, steven. 1989. learnability and cognition: the acquisition of argument structure. the mit press. portner, paul. 2004. the semantics of imperatives within a theory of clause types. in semantics and linguistic theory, vol. 14, 235–252. portner, paul. 2007. imperatives and modals. natural language semantics 15(4). 351–383. https://doi.org/10.1007/s11050-007-9022-y. rakoczy, hannes & michael tomasello. 2009. done wrong or said wrong? young children understand the normative directions of fit of different speech acts. cognition 113(2). 205–212. https://doi.org/10.1016/j.cognition.2009.07.013. rett, jessica. 2011. exclamatives, degrees and speech acts. linguistics and philosophy 34(5). 411–442. https://doi.org/10.1007/s10988-011-9103-8. roberts, craige. 1996. towards an integrated formal theory of pragmatics. osu working papers in linguistics 49. roberts, craige. 2004. context in dynamic interpretation. in laurence r. horn & gregory ward (eds.), the handbook of pragmatics, 155–196. oxford and malden, ma: blackwell. roberts, craige. 2018. speech acts in discourse context. in new work on speech acts, oxford university press oxford. https://doi.org/10.1093/oso/9780198738831.003.0012. sadock, jerrold m & arnold m zwicky. 1985. speech act distinctions in syntax. language typology and syntactic description 1. 155–196. saffran, jenny r, richard n aslin & elissa l newport. 1996. statistical learning by 8-month-old infants. science 274(5294). 1926–1928. https://doi.org/10.1126/science.274.5294.1926. searle, john r. 1975. indirect speech acts. in peter cole & j. morgan (eds.), speech acts (syntax and semantics 3), 59–82. new york: academic press. stalnaker, robert c. 1978. assertion. in pragmatics, 315–332. brill. starr, william b. 2014. mood, force and truth. protosociology 31. 160–181. https://doi.org/10.5840/protosociology20143113. starr, william b. 2020. a preference semantics for imperatives. semantics and pragmatics 13. 6. http://dx.doi.org/10.3765/sp.13.6. woodward, amanda l. 1998. infants selectively encode the goal object of an actor’s reach. cognition 69(1). 1–34. https://doi.org/10.1016/s0010-0277(98)00058-4. proceedings of elm 1: 284-297, 2021 anissa zaitsu, jad wehbe, valentine hacquard and jeff lidz: clause types and speech acts in speech to children. 297 https://doi.org/10.3765/elm https://www.elm-conference.net/ simulating semantic change: a methodological note remus gergel, martin kopf-giammanco, & maike puhl* abstract. the current work discusses the human diachronic simulation paradigm (hudspa), a method to experimentally probe into historical meaning change set up to (i) scan for configurations similar to attested alterations of meaning but in (typically, but not necessarily, related) languages or varieties which did not actualize the change(s) under investigations; (ii) measure the reactions of native speakers in order to ascertain the verisimilitude as well as the particular semantic and pragmatic properties of the items scrutinized. specifically, the present paper discusses the relative propensity of a particularizer (german eben) to be interpreted with comparatively high confidence as a scalar additive particle such as even and of a concessive item like english though to be interpreted similar to a modal particle along the lines of german doch. keywords. diachronic semantics; additives; modal particles; experimental semantics; experimental pragmatics 1. introduction. diachronic and fieldwork semantics both model natural language variation. however, their standard methods of empirical verification vary considerably. sometimes, they are even viewed as not (yet) fully compatible. for instance, deal (2020) considers variable-force modals in synchronic and diachronic semantics and raises questions about diachronic conclusions (e.g. when variable-force semantics is suggested based on a sample of 72 old english examples). our present goal is not to engage with particulars of old english modality (cf. cournane 2017, gergel 2016, 2017, yanovich 2006, solt & umbach 2019 for broader discussions of natural-language modality from different perspectives with relevance to historical studies). but the more general point raised by deal holds and needs to be addressed systematically. diachronic semantics is constrained in multiple respects and to some serious extent this appears to be due to its intrinsic nature, which seems to run counter to methods of inquiry used in, say, modern cross-linguistic semantics. (ultimately, we do not think that such a putative incompatibility is a necessary conclusion, as we will see.) hence, regardless of the origins of possible empirical dissonances and difficulties in diachronic semantics, a continuous refinement of the empirical methods that are used in this branch seems to us to be imperative and useful.1 *we thank the organizers, reviewers, and audience of elm 1 for providing a fruitful and encouraging forum of discussion for our ideas. authors: remus gergel, universität des saarlandes (remus.gergel@uni-saarland.de) & martin kopf-giammanco, universität des saarlandes (martin.kopf@uni-saarland.de) & maike puhl, universität des saarlandes (maike.puhl@uni-saarland.de) 1the value of diachronic studies themselves is not at issue here or in any of the studies mentioned and it should be evident, but we offer in oversimplified fashion an argument to readers less familiar with the field. diachronic data is transitional data. that is, such data not only provides descriptions of, e.g. a well-studied language (variety) a and an exotic language (variety) b based on a parametric or otherwise postulated discriminating modelling criterion, but it is concerned with an old language (variety) o and a new language (variety) n that has developed over time from o. typically this requires a more constrained theoretical modelling, as an explanation for how to plausibly and factually reach from o to n is by default required. proceedings of elm 1: 184-196, 2021 c©2021 remus gergel, martin kopf-giammanco, and maike puhl published by the lsa with permission of the author(s) under a cc by license. 184 https://doi.org/10.3765/elm https://www.elm-conference.net/ we wish to emphasize that diachronic semantics has not been static over recent years, but it has made considerable progress; see e.g. deo (2015) for an overview, to which we only add a few relevant points before narrowing down further to our current point. importantly, clear theoretical programs exist (cf. e.g. fintel 1995, eckardt 2006) and likewise specific corpus studies, for instance even to make sense of conflicting analyses that could not be previously solved synchronically (beck & gergel 2015, gergel & beck 2015) as well as studies that connect somewhat broader typological views with methods as refined as electrophysiological measurements (cf. zhang et al. 2018). when not all the data relevant for semantic change is available, researchers have moreover capitalized on the interfaces of the semantic component, e.g. with structure or with pragmatic processes (cf. gianollo 2018, traugott 2006, to name just one example for each possibility in this case as well). furthermore, the lack of extensive corpus data for low-frequency phenomena can sometimes be partially compensated when more methods are amassed (see gergel & kopf-giammanco forth. for an example and discussion). and nonetheless, when it comes to the necessary details of virtually any semantic inquiry, there are impasses that appear to be often insurmountable from the perspective of historical linguistics if one takes fieldwork variationist semantics as a term of comparison. it is not just the lack of negative data that poses a well-known problem. receiving graded judgments in appropriate and detailed contexts from native speakers to test the validity of both proposed trajectories and causal chains in actual changes can be just as critical as the lack of such data can be frustrating. such issues motivate a more general question: how can a semantic path of change (and especially under the inclusion of typologically less well-trodden ones) be established given the impossibility of eliciting contextualized judgments, of receiving comments from interviewed speakers, etc.? some further venues are conceivable to compensate such deficits and offer partial answers to the question raised. for example, a possible solution to the main problem is to analyze current changes in progress (e.g. d’arcy 2007), ideally such that they resemble changes that have occurred in the past. however, there are by far not enough detectable changes in progress to match the numerous interesting meanings that have arisen historically. therefore, we will support the use of experimental semantics as a bridge to cross the gap between semantic fieldwork and diachrony drawing from other cases of semantic development with disadvantaged extraction of speaker intuitions, namely the earliest stages of language acquisition (cf. e.g. gleitman et al. 2005). specifically, we discuss in this paper two experiments to support the hypothesis in (1) (gergel 2020: 13): (1) human diachronic simulation paradigm (hudspa) humans confronted with new meaning-form pairings modeled after an attested semantic change will react similarly when they are placed in conditions that resemble those of the actual change (e.g. via a cognate that is similar but did not undergo the transformation investigated). the idea is to confront speakers with meanings that happened in a related language or variety but not in their own and to compare them to meanings that did not develop in either language or variety. following hudspa, we hypothesize that the meanings that were targeted in the actuated change will perform better than other meanings that could have developed from the same semantic domain. we thus seek to replicate relevant parts of historical processes under testable conditions. proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 185 https://doi.org/10.3765/elm https://www.elm-conference.net/ in other words, we target and sort out suboptimal (“ungrammatical”) but relevant judgements with a potential for change in pertinent environments. speakers in historical change situations also end up with choosing form-meaning pairings that are originally nontransparent on the meaning that will later conventionalize. but such choices are less random than they may appear. the hypothesis we support in this contribution is precisely that they can be replicated to a degree in lab conditions. in section 2, we will present two experiments in which we confronted german speakers with contexts in which the meaning of english even is conveyed, by utilizing a cognate (german eben), and english speakers with contexts in which a modal-particle meaning modeled after german doch is utilized for the word though. section 3 provides a detailed discussion of the results and section 4 concludes and provides an outlook of how hudspa experiments can be refined further. 2. two experiments. our first experiment targets the development of english even, simulated from the perspective of german eben. german eben did not develop a meaning such as (modern) english even. german only uses noncognates of eben for additives of improbability. the meaning of eben can be approximated to the meaning of even in contexts such as an even surface. our second experiment targets the german discourse particle doch through the prism of english though. we paid attention to syntax, e.g. by using final though in view of relevant factors (van kemenade 2019). similar to above, though did not develop a presuppositional meaning as doch (grosz 2014), regardless of syntax. in both experiments we used two cues to activate speakers to such readings: one is context to clarify the intended meaning; the other is the instruction to treat the examples as spoken by some non-mainstream german (and english) community and to grade the naturalness of the examples encountered w.r.t. to the context given. from an earlier study, we had confirmation that speakers can reliably assign meanings in rich contexts to sentences which they find otherwise unacceptable (cf. e.g. gergel 2020, gergel & kopf-giammanco forth. for discussions). 2.1. eben manipulated as english even. in this experiment, a questionnaire with 12 target items and 13 filler items was used. the target items consisted of 3 item sets with each set consisting of 4 items and respectively licensing readings of sogar (‘even’), nur (‘only’), and auch (‘too/also’). in place of sogar, nur, and auch, the items featured eben –– cf. fig. 1 for an example item. all items consisted of a context description (letztes wochenende hatten wie eine große party; ‘last weekend we had a big party’; fig. 1 top) and a target sentence (eben maria, die sonst immer zuhause bleibt, ist gekommen; ‘eben mary who usually stays at home showed up’; boldfaced in fig. 1) as well as a comment section. subjects were asked to rate the target sentences based on a 7-point scale ranging from ‘fully acceptable in context’ (7 pts) to ‘not at all acceptable in the context’ (1 pt). in the comment section, subjects were encouraged to suggest improvements should they find certain expressions odd. we collected data from 71 subjects, all of them undergraduate students in the english department of saarland university, yielding 810 observations (excluding 42 missing values from the ratings). we excluded non-native german speakers. additionally, we manually categorized the comments provided by the subjects as to their suggestions for improving the target sentences. the criterion here was that the subjects suggested replacing eben with sogar/nur/auch. if subjects suggested supplementing eben with sogar/nur/auch, commented on an unrelated issue (or did not provide a comment at all) their rating was not considered for this analysis. this criterion was cruproceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 186 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: example item experiment 1; eben manipulated for even/sogar cial because eben can be used in connection with sogar/nur/auch but is interpreted as a discourse particle rather than with the targeted meaning (cf. repp 2013). based on this manual categorization, we had 199 observations (53 for the sogar condition, 94 for nur, 52 for auch) for further analysis. in descriptive terms, the three conditions were rated as in table 1. sogar-‘even’ nur-‘only’ auch-‘also/too’ mean 5.17 4.34 4.62 median 6 5 5 sd 1.46 1.6 1.83 table 1: mean and median ratings of experiment 1 for statistical analysis, we relied on the r software (r core team 2019) and the lme4-package (bates et al. 2015) for r. in a first step, we transformed the ratings into norm scores2 and fit the data into a random slope model with ‘normscore’ as a function of ‘condition’ (i.e. the 3 levels: sogar, nur, auch), allowing for different slopes per subject: normscore ∼ condition + (1 + condition | subject) (cf. bates et al. 2015, r core team 2019). the estimate for the sogar(‘even’)-level is 0.222 and the slope for the nur(‘only’)-level is -0.561, for auch(‘also/too’) -0.382. in a second step and in order to obtain a p-value, we conducted a likelihood ratio test, pitching the full model against a null model (i.e. without the factor of interest, ‘condition’). the three levels of the factor condition affected the transformed ratings (χ2 (2) = 13.221, p=.0013) lowering them by 0.561 for the nur-level and by 0.382 for the auch-level. this comparison suggests that the variability in the data collected is not random but can be explained by the three levels of the experiment. 2.2. though manipulated as german doch. in the second experiment, manipulating final though as doch, an online questionnaire was used with 12 target items (joined by 14 fillers) with 4 target items per condition, where the respective readings approximated three different types of particles: doch, ja, wohl (cf. zimmermann 2011 for an overview of the untranslatable material and puhl & gergel forth. for a discussion on the meaning contribution of final though). as an approximation, the modal particle ja marks an utterance p as uncontroversial because 2normscore transformation was performed in order to account for inter-subject-variability and achieve more normally distributed values. proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 187 https://doi.org/10.3765/elm https://www.elm-conference.net/ p is already in the common ground (cg). according to repp (2013), ja fulfils a retrieval function, meaning that the speaker’s use of ja instructs the addressee to retrieve a proposition p from the cg (repp 2013). this proposition p is not under consideration at the time of utterance, meaning that it is not entailed or implicated by the immediately preceding utterance (repp 2013). consider (2): (2) ich i gehe go heute today nicht not spazieren. walk. es it regnet rains ja. prt. ‘i’m not going for a walk today. it is raining, you know.’ in (2), the weather is not entailed or implicated by the speaker’s decision not to go for a walk. the hearer is assumed to be aware of the weather it is part of the cg but it is not being considered at the time of utterance. the particle doch is similar to ja in that it also instructs the hearer to retrieve a proposition from cg. the difference between ja and doch is that doch also signals a contrast. following repp (2013), this contrast lies between the proposition p in doch(p) and a proposition q (= ¬p) in the cg. (3) ich i gehe go heute today nicht not spazieren. walk. # # es it regnet rains doch. prt. ‘i’m not going for a walk today. # but it is raining, you know.’ (4) a: i’m going for a walk now. b: es it regnet rains doch. prt. ‘but it’s raining, as you know’ the use of doch in (3) is infelicitous because there is no contrast between not going for a walk and bad weather3. in (4), the use of doch is felicitous. it is assumed that both speakers a and b are aware of the weather. doch signals this and instructs a to retrieve this information from cg. a’s decision to go for a walk is at odds with the fact that people tend to go for walks in good weather4. this contrast between bad weather and going for a walk licenses the use of doch and, at the same time, makes the use of ja infelicitous, see (5). (5) a: i’m going for a walk now. 3felicitous readings of (3) are possible depending on intonation and an assumed hearer. if “es regnet doch” is the answer to an implied “why not?”, (3) is acceptable because it highlights that the hearer is not currently considering the weather. however, if contrasted with (2) and keeping intonation constant, (3) is infelicitous. 4note that if a has a habit of going for walks in the rain and b knows this, a dialogue such as (i) is perfectly acceptable. (i) a: i’m not going for a walk today. b: es it regnet rains doch. prt. ‘but it’s raining, as you know’ b’s utterance can be paraphrased as: “why are you not going for a walk? it’s raining and you like walking in the rain”. proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 188 https://doi.org/10.3765/elm https://www.elm-conference.net/ b: # # es it regnet rains ja. prt. ‘# it’s raining, you know.’ the modal particle wohl has no overlap with either ja or doch. it is an epistemic marker signaling that the speaker is not fully committed to the utterance, but merely assumes the utterance to be true (e.g. zimmermann 2004, 2011), as in (6). (6) a knows that b wanted to go for a walk. a: es it regnet. rains. dann then geht goes sie she wohl prt nicht not spazieren. walk. ‘it’s raining. i assume she’s not going for a walk’ in (6), a does not know whether b ended up going for a walk despite the weather but can only assume that this is not the case. trying to reproduce felicitous readings of ja, doch and wohl required the use of slightly longer and only dialogic contexts, compared to the first experiment. given that the particles do not have counterparts in english, the experiment included two tasks, the first one consisting of a training section and asking if the meaning from the contextual clues was understood. the answer to this task was given through a slider ranging from 1 (‘very hard to understand’) to 101 (‘very easy to understand’). subjects were also asked to provide a paraphrase of what they assume is meant by this sentence. given that the language in which this experiment was conducted, english, lacks the particles, it could not be expected to have the same precision in the additional comments as in experiment 1. the second task was a forced-choice yes/no slider to test if the item was actually understood, i.e. whether final though conveyed the intended meanings of doch, ja and wohl, respectively. see figure 2 for an example item of the second task. figure 2: example item experiment 2; though manipulated for doch 40 native speakers of english participated in this experiment, but due to the inclusion of attention-testing fillers, only 36 were considered. these attention-testing fillers included specific instructions in the context sentence about where to move the slider regardless of how good the supposed target sentence was, e.g. “move the slider all the way to the left”. see table 2 for the descriptive statistics of this experiment. while the sentences seemed easy to understand (task 1), proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 189 https://doi.org/10.3765/elm https://www.elm-conference.net/ the intended meanings were not captured reliably (task 2). task 1 doch ja wohl mean 95.40 92.47 92.15 sd 10.48 18.53 17.53 task 2 mean 84.45 59.94 23.73 sd 25.33 38.34 34.07 table 2: mean ratings of experiment 2 comments were analyzed by assigning each comment a category which best fits the content of the comment. 14 categories in total included paraphrase (target sentence without though), reminder (doch), common ground (ja and doch) and assumption (wohl). other important categories are reason/explanation (most common) and concessive (use of though in pde). for doch, almost 40% of the comments fell into the categories reminder and common ground, which closely resemble the meaning of the particle. for ja, the most common category was explanation/reason (38%), which does not capture the meaning of the particle ja. better suited categories, such as common ground or knowledge, add up to less than 2%. the most common category of comments for wohl was also explanation/reason with 71%. again, this does not capture the intended meaning of wohl. assumption, which best fits wohl, received 18%. we conducted exact wilcoxon signed rank tests for each pair (doch-ja, doch-wohl, ja-wohl for tasks 1 and 2, and doch1-doch2, ja1-ja2, and wohl1-wohl2). in task 1, doch was rated significantly higher than wohl (p = 0.002) but there are no significant differences between doch and ja (p = 0.093), and ja and wohl (p = 0.740). doch readings were rated significantly higher in task 2 than ja and wohl, and ja was rated significantly higher than wohl (p < 0.001 in all three cases). all three target readings showed significant differences between tasks 1 and 2 (doch: p = 0.004, ja: p < 0.001, wohl: p < 0.001). 3. discussion. the experiments show that the meanings of the cognates were interpreted more appropriately than competitors. both the doch meaning of though and the even meaning of eben were captured significantly more reliably than the meanings of their competitors. the discourse particle meaning of doch seems calculable from the relationship between the currently available concessive component of final though, which was reflected in the comments, which is close in meaning to the presupposition of contrast in doch. both doch and though are possible in concessive contexts, see (7) and (8). (7) a: i’m going for a walk. b: it’s raining, though. (8) a: i’m going for a walk. b: es it regnet rains doch. prt. ‘but it’s raining, as you know.’ in both (7) and (8), the contrast between bad weather and going for a walk is signaled by though and doch, respectively. proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 190 https://doi.org/10.3765/elm https://www.elm-conference.net/ nevertheless, with only 40% of the comments falling into the categories of reminder and common ground, which most accurately describe the meaning and usage of the mp doch, the differences between the cognates though and doch remain apparent. while both doch and though convey the notion of contrast, they differ in their conditions of use due a key point of difference: doch includes a retrieval function while though does not. doch can be used to signal a reminder. consider (9) and (10). (9) a: it’s too bad that tim isn’t coming to the party. b: er he kommt comes doch. prt. ‘he’s coming, remember?’ (10) a: it’s too bad that tim isn’t coming to the party. b: he’s coming, though. both (9) and (10) are felicitous. however, in (9), the implicature arises that a should have known (or did know at some point) that tim was coming to the party. in (10), no such implicature arises. this implicature is easily cancellable, as in (11). (11) a: it’s too bad that tim isn’t coming to the party. b: er he kommt comes doch. prt. ‘he’s coming, remember?’ a: really? i didn’t know that. in the experiment, this reminder function of doch was reinforced. participants appear to have identified the notion of contrast that is present in doch and though but seem to have struggled with the retrieval component of doch. the additive case may seem more surprising. however, if we consider that german eben can have e.g. a meaning similar to what traugott (2006) identifies as a particularizing focus modifier reading (pfm; for early english even), as in (12), then we can explain the significantly higher acceptability ratings for the items where eben was manipulated for even. traugott describes such a reading of even as precursor (stage ii of a 3-stage development) towards the modern one. beaver & clark (2008) characterize particularizers as typically non-scalar focus operators. they propose (i) that their use indicates that a speaker has provided an indication of being in a position to answer the (possibly implicit) current question (cq) and (ii) that specific particularizers possibly provide additional information. along these lines, german ebenpfm as in (12) could be viewed as – aside from beaver & clark’s (i) – indicating that, among the alternatives, there is exactly one possible candidate to answer the cq (who did peter meet?) and that the focused individual is salient. for scalar additives, beaver & clark note that they state that the strongest true answer to the cq is weaker than expected – in other words, the prejacent is the most improbable proposition from a set of alternatives.5 given the possible availability of eben as in (12) in the subjects’ grammars, they might have had an easier time accommodating a scale of (im)probability rather than for items where eben was manipulated for also/too and only. specifically, the salience of the entity singled 5strength is by default characterized by entailment; however, see beaver & clark (2008; 254ff) for details. proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 191 https://doi.org/10.3765/elm https://www.elm-conference.net/ out by ebenpfm might have been responsible for participants’ higher acceptability for its use in even-contexts. (we surmise, however, that in appropriate settings, salience might also offer a considerable bias for only-readings, even if the eben-readings have been found significantly the most acceptable ones in our experiment. we will return to this below.) (12) peter peter hat has letzte last woche week im in.the krankenhaus hospital maria maria kennengelernt. gotten to know eben prt [(diese) this maria]f maria hat has er he heute today beim at einkaufen shopping getroffen. met ‘last week, peter and maria met at the hospital for the first time. today at the grocery story, he ran into exactly that maria.” in turning to eben manipulated for only, we want to point out that the only-meaning can be associated with the use of eben only in varieties of austrian german, cf. (13). subjects evaluated in the experiment reported here were all native speakers of federal german. (13) ich i habe have so so viel much zu to tun, do, aber but ich i habe have eben prt zwei two hände. hands. ‘i have so much to do, but i have only two hands.’ (only varieties of austrian german; adapted for standard orthography) as noted above, the only-sentences were “guessed” correctly more often (94 vs. 53 times for evenand 52 times for also/too-sentences; via the improvement task). nonetheless, recall that acceptability was rated significantly highest and most appropriately on the even readings. we think, the set of facts we have so far is less recalcitrant than it may appear at first glance not only from an intuitive perspective (as salience could possibly be made to play an important role to different degrees in both even and only readings) or from a variationist perspective (as some dialects did in fact develop some only readings of the particularizer). if we take a step back and consider the broader picture from the point of view of semantic theory, it is also not too surprising that the two types of readings may compete in a close race historically (e.g. when contexts of change are indeterminate or biased one way or another). beaver & clark suggest that scalar additives (as even) and exclusives (e.g. only) are close-by in a certain sense and represent pragmatic opposites. while scalar additives amount to stating that the strongest true answer is stronger than expected, for exclusives the strongest true answer to the cq is weaker than expected. in (14), the prejacent mary and phil came to the party is the strongest true answer to the cq who came to the party?. the upper bound placed on the strength of possible answers to the cq by only is where its truth conditional impact originates: any possible answer with more individuals than mary and/or phil showing up to the party is not true. it seems that contextual clues pertaining to truth conditions in the experimental items provided participants in the particular set-up with more solid ground for identifying the intended meaning of eben6. 6in designing our experimental items, we heavily relied on the upper bound that only triggers by using e.g. examples as in (i) (translated from g to e for ease of representation): (i) context: the seminar this past term was a lot of work. students had to submit practice sheets every week. although they tried their best, just before the end of term it was too much for everybody: proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 192 https://doi.org/10.3765/elm https://www.elm-conference.net/ (14) only mary and phil came to the party. (15) even mary came to the party. even (cf. (15)) does not have an upper bound that causes a truth conditional effect (any stronger alternatives are not necessarily excluded) but it conversely pushes the upper bound of what is expected to be the strongest true proposition. it is even’s capacity to surprise which we suspect to have given it an edge over only in terms of acceptability with their respective meanings forced onto eben –– especially, as noted above, in face of eben’s particularizer function and its condition that the focus element in its clause be salient. this last point also supports the core hypothesis in hudspa that when appropriately contextualized, originally nontransparent form-meaning pairs can be accommodated and conventionalized along an actual trajectory of semantic change. 4. summary and outlook. to sum up, hudspa at this point shows convergence towards the actually developed meaning if the speakers’ grammars are properly factored out. this is a minimal but crucial result towards more refined investigations of change. while, for instance, several currently popular game-theoretic approaches have a similar goal of simulating paths of change on the basis of rational tools, they do so at times by stipulating (often rather abstract) costs and benefits, so that in principle nearly any course could be attained. hudspa, by contrast, constrains the course of change appropriately, by using as its primary sources solely natural-language intuitions, which can further be probed into experimentally and theoretically. from a broader perspective, we think hudspa should not be all that surprising. it generalizes a certain perspective on uniformitarianism (cf. walkden 2019 for a rich historical discussion even if without a semantics excursus) and the well-known idea from several branches of linguistics including sociolinguistics and language acquisition that the present is worth considering also to explain aspects of the past. what we take to be just as worthwhile is a starting attempt to raise such questions in controlled experimental environments for the area of semantic change. while the results of the experiments above seem to support hudspa, they can only be regarded as initial findings. refining experimental design based on hudspa is the goal of follow-up research. we mention here only a few ways how we envisage the paradigm can be refined in future work. a first extension entails the incorporation of training tasks, which could be additionally tested (also via interactive experimental design, targeting different types of memory storage, etc.). a second way to refine the insight obtained is to test if results differ depending on whether or not participants received the instruction that they are encountering a non-standard variety. to some extent, this could be viewed as paralleling a putative dichotomy of contact-induced, i.e. external vs. endemic changes. but notice that from the perspectives of speakers who have not yet adopted a given semantic change, contact with progressive speakers with regards to the change in question is in practice almost always the case (even when they belong to the same linguistic community in other respects). a third controlling step would be to rule out additional possible biases ranging from less obvious lexical semantic facts to relevant phonological biases. the clearest cut in this area would target: out of all students, eben mary submitted her homework on time. proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 193 https://doi.org/10.3765/elm https://www.elm-conference.net/ seem to be offered by the use of nonce words, but in such a case too, possible associations with relevant words known by participants would have to be controlled. (e.g. if a nonce word is more similar to an existing relevant word than competitors, then it could still have a starting advantage.) the usage of nonce words would clearly reduce the ”etymological burden” in design, but notice that this by and of itself does not automatically offer an improved insight. speakers in actual change situations quite often in fact take the previous meaning as a starting package and build on it through interactive processes. however, a controlled design could target nonce words that have been introduced (and crucially: trained) in very specific ways, so that only those features will figure prominently that are relevant to the experimental task. last but not least, we have only illustrated a minimal amount of variation in the methods used for hudspa for practical reasons; there is naturally no a priori reason to constrain either the technical battery of methods or the range of applications (a quick extension could be, for instance, to look beyond naturalistic l1 semantic changes and also incorporate the potential of different types of bilingual or l2 extensions.) the major restriction remains, just as much as the potential that we see, that a close investigation of the actual critical contexts of change (as opposed to possibly too broad generalizations) may be one of the safest ways to plausibly simulate semantic change. references bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67(1). 1–48. 10.18637/jss.v067.i01. beaver, david & brady clark. 2008. sense and sensitivity. how focus determines meaning. chichester, uk: wiley-blackwell. beck, sigrid & remus gergel. 2015. the diachronic semantics of english again. natural language semantics 23,3. 157–203. cournane, ailı́s. 2017. in defence of the child innovator. in éric mathieu & robert truswell (eds.), micro change and macro change in diachronic syntax, 10–24. oxford: oup. d’arcy, alexandra. 2007. like and language ideology: disentangling fact from fiction. american speech 82. 386–419. deal, amy-rose. 2020. comments on diachronic formal semantics (as compared to formal semantic fieldword). position paper given at the workshop formal approaches to grammaticalization. lsa annual meeting, new orleans, january 5, 2020. deo, ashwini. 2015. diachronic semantics. annual review of linguistics 1. 179–197. eckardt, r. 2006. meaning change in grammaticalization. an enquiry into semantic reanalysis. oxford: oup. fintel, kai von. 1995. the formal semantics of grammaticalization. in proceedings of nels 25 2: papers from the workshops on language acquisition & language change, 175–189. umass, amherst. gergel, remus. 2016. modality and gradation: comparing the sequel of developments in ‘rather’ and ‘eher’. in elly van gelderen (ed.), the linguistic cycle continued, 319–350. amsterdam/philadelphia: john benjamins. gergel, remus. 2017. dimensions of variation in old english modals. in ana arregui, maria luisa proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 194 https://doi.org/10.3765/elm https://www.elm-conference.net/ rivero & andres salanova (eds.), modality across syntactic categories, 179–207. oxford: oxford university press. gergel, remus. 2020. sich ausgehen: actuality entailments and further notes from the perspective of an austrian german motion verb construction. in proceedings of the lsa workshop ”formal approaches to grammaticalization”, vol. 5(2), 5–15. new orleans: linguistic society of america. gergel, remus & sigrid beck. 2015. early modern english again: a corpus study and semantic analysis. english language and linguistics 19,1. 27–47. gergel, remus & martin kopf-giammanco. forth. ‘sich ausgehen’: on modalizing goconstructions in austrian german. canadian journal of linguistics . gianollo, chiara. 2018. indefinites between latin and romance. oxford: oup. gleitman, lila, kimberly cassidy, rebecca nappa, anna papafragou & john trueswell. 2005. hard words. language learning and development 1. 23–64. grosz, patrick. 2014. german doch: an element that triggers a contrast presupposition. in proceedings of the 46th annual meeting of the chicago linguistic society, 163–177. puhl, maike & remus gergel. forth. final though. in remus gergel, ingo reich & augustin speyer (eds.), particles in german, english and beyond, amsterdam: benjamins. r core team. 2019. r: a language and environment for statistical computing. r foundation for statistical computing vienna, austria. https://www.r-project.org/. repp, sophie. 2013. common ground management: modal particles, illocutionary negation, and verum. in daniel gutzmann & hans-martin gärtner (eds.), beyond expressives: explorations in use-conditional meaning, 231–274. leiden, the netherlands: brill. solt, stephanie & carla umbach. 2019. comparison via eher. paper presented at sinn & bedeutung 24, university of osnabrück. traugott, elizabeth. 2006. the semantic development of scalar focus modifiers. in ans van kemenade & bettelou los (eds.), the handbook of the history of english handbooks in linguistics, 335–359. malden et al.: blackwell. van kemenade, ans. 2019. discourse particle then in english: clause structure and discourse organisation. paper presented at particles in german, english, and beyond, saarland university. walkden, george. 2019. the many faces of uniformitarianism in linguistics. glossa: a journal of general linguistics 4(1). 52. 1–17. yanovich, igor. 2006. old english *motan, variable-force modality, and the presupposition of inevitable actualization. language 92(3). 489–521. zhang, muye, maria mercedes piñango & ashwini deo. 2018. real-time roots of meaning change: electrophysiology reveals the contextual-modulation processing basis of synchronic variation in the location-possession domain. in c. kalish, m rau, j. zhu & t.t. rogers (eds.), proceedings of the 40th annual conference of the cognitive science society, 2783–2788. austin, tx: cognitive science society. zimmermann, malte. 2004. zum “wohl”: diskurspartikeln als satztypmodifikatoren. in linguistische berichte 199, 253–286. hamburg: helmut buske verlag. zimmermann, malte. 2011. discourse particles. in paul portner, claudia maienborn & klaus von heusinger (eds.), semantics handbücher zur sprachund kommunikationswissenschaft, proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 195 https://doi.org/10.3765/elm https://www.elm-conference.net/ hsk 33.2, 2011–2038. berlin: mouton de gruyter. proceedings of elm 1: 184-196, 2021 remus gergel, martin kopf-giammanco, and maike puhl: simulating semantic change: a methodological note. 196 https://doi.org/10.3765/elm https://www.elm-conference.net/ tattoos as a window onto cross-linguistic differences in scalar implicature danielle dionne & elizabeth coppock* abstract. this paper addresses the question of how to predict which alternatives are active in scalar implicature calculation, and the nature of this activation. it has been observed that finger implicates ‘not thumb’, and a manner-based explanation for this has been proposed, predicting that if english had the simplex latin word pollex meaning ‘thumb or big toe’, then finger would cease to have the implicature ‘not thumb’ that it has. it has also been suggested that this hypothetical pollex would have to be sufficiently colloquial in order to figure in scalar implicature calculation. this paper makes this thought experiment into a real one by using a language that behaves in exactly this way: spanish has pulgar ‘thumb’ (< pollex), a non-colloquial form. we first use a fill-in-the-blank production task with both english and spanish speakers to gauge the likelihood with which a speaker will produce a given form as a way of describing a given digit. production frequency does not perfectly track complexity, so we can then ask whether comprehension follows production frequency or complexity. we do so using a forced choice comprehension task, which reveals cross-linguistic differences in comprehension tracking production probabilities. a comparison between two rsa models – one in which the speaker perfectly replicates our production data and a standard one in which the speaker chooses based on a standard cost/accuracy tradeoff – illustrates the fact that comprehension is much more closely tied to production probability than to the mere existence of sufficiently simple alternatives. keywords. scalar implicature; manner implicature; hyponymy, cross-linguistic differences; rsa; computational modelling 1. introduction. suppose you heard the following sentence: (1) she has a tattoo on her finger. would you think the tattoo was on the thumb or the ring finger? if you are like most of the participants in our english comprehension study, you will think the ring finger is more likely. the thumb is generally considered a type of finger; people generally agree that we have 10 fingers. so it is arguably not the semantics of finger that determines this preference; rather, there is a scalar implicature from finger to ‘not thumb’. in gricean terms, the pragmatic reasoning might run as follows, ‘why didn’t she choose thumb? it would have been equally short (manner), more informative (quantity), and just as relevant (relevance). maybe she didn’t believe it (quality).’ horn (2000) observes that the relationship between thumb & finger is not parallel to the relationship between big toe & toe: (2) a. i hurt my finger. ↝ i did not hurt my thumb. b. i hurt my toe.   i did not hurt my big toe. *we are grateful to the audiences at lsa 2020 and elm 2021 for feedback on this work. authors: danielle dionne, boston university (ddionne@bu.edu) & elizabeth coppock, boston university (ecoppock@bu.edu). proceedings of elm 1: 147-158, 2021 c©2021 danielle dionne and elizabeth coppock published by the lsa with permission of the author(s) under a cc by license. 147 https://doi.org/10.3765/elm https://www.elm-conference.net/ horn (2000) concludes that although thumb acts as an alternative to finger for the purposes of scalar implicature, big toe does not act as an alternative for toe (p. 308). he explains this in terms of manner: big toe is longer than toe, and therefore not a good alternative. horn (2000, p. 308) writes: “we would predict that if the colloquial language replaced its thumb with the polymorphous pollex (the latin and scientific english term for both ‘thumb’ and ‘big toe’), the asymmetry [between finger and toe] would instantly vanish”. geurts (2011) zeroes in on horn’s strategic use of the term “colloquial”, writing: “it is important to note, however, that the adjective ‘colloquial’ is doing real work in this statement. it is not enough for an alternative word to be in the language; it has to be sufficiently salient, as well: if the word ‘thumb’ was rarely used, then presumably the asymmetry between [finger and toe] would vanish too” (p. 122). that is, the prediction is really that if a stronger utterance is present in the language and it is sufficiently salient, a scalar implicature will arise when the weaker form is used. as it turns out, there is a language in which exactly that situation arises: spanish. spanish contains a word for ‘thumb’, namely pulgar—the spanish descendant of latin pollex—but it is less frequently used, and less colloquial. as we will confirm in production studies, there is a great deal of variation in how the thumb is referred to in spanish. pulgar does not differ from thumb in complexity, but it does differ in how prevalent it is in the language, and how salient it is to the speaker as an alternative. furthermore, pulgar ‘thumb’ is not specific to the hand, just like latin pollex. if geurts (2011) is right, the asymmetry between finger and toe that exists in english is predicted to be absent in spanish, and there should be no implicature from dedo ‘finger’ to ‘not thumb’ or from dedo del pie ‘toe’ to ‘not big toe’. the four studies reported here test these predictions. we first conducted production studies each in english and spanish to gauge the salience of available alternatives. we use a fill-in-theblank production task with both english and spanish speakers to gauge the likelihood with which a speaker will produce a given form as a way of describing a given digit. we find that production frequency does not perfectly track complexity, so we can then ask whether comprehension follows production frequency or complexity. we do so using a forced choice comprehension task, which reveals cross-linguistic differences in comprehension that tracks production probabilities. finally, we carry out a comparison between two rsa models, one in which the speaker perfectly replicates our production data and a standard one in which the speaker chooses based on a standard cost/accuracy trade-off. comparing both of these models to our comprehension data leads us to conclude that the activation of alternatives for the purpose of scalar implicature calculation is much more closely tied to production probability than to the mere existence of sufficiently simple alternatives. 2. production studies. we conducted two production studies: one in english and one in spanish. both tasks contained images of body parts with tattoos (see figure 1). 2.1. participants. all participants were recruited on prolific. all studies involved different groups of participants. in the english production study, all participants were self-reported monolingual native english speakers who were born and currently live in the united states. in the spanish production study, all participants were self-reported monolingual native spanish speakers who were born and currently live in mexico. there were 24 american english speakers and 23 mexican spanish speakers in the production studies. proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 148 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: stimulus items for production and comprehension tasks 2.2. materials. participants completed a task in which they were asked to look at a series of pictures. all of the pictures were body parts with a tattoo on them. the tattoos served as an indicator of which digit or body part the speaker was talking about. the target items showed photos of a tattoo on the thumb, ring finger, pinky, big toe, fourth toe and pinky toe. there were six filler items (all different from each other): two photos of tattoos on a leg, two photos of tattoos on an arm, and two photos of tattoos on the back. 2.3. procedure. participants were shown a series of images, one by one. with each image, they were asked to fill in the blank of the sentence: she has a tattoo on , or its translational equivalent, in the case of spanish. the order of images was randomized. all participants were presented with all six target items and all six filler items. 2.4. normalizing production results. after the data collection process was completed, all responses for the production study were normalized by hand. this included removing additional words such as “left” or “right” (e.g. “right pinky” became “pinky”). directional terms (“left”), initial articles, and other non-essential words were stripped away so that all that was remaining was the word or phrase that was used to refer to the digit itself. this removed excess noise from the data and allowed us to group responses together that were essentially identical in form – dedo de la mano (‘digit of the hand’ finger) vs. dedo (‘digit’), for example. additionally, responses were coded for specificity — 1 for specific words/phrases that could refer to only one digit (e.g. “thumb” or “pulgar”) and 0 for non-specific words/phrases that could refer to more than one digit (e.g. “finger” or “dedo de la mano”). 2.5. results. for the thumb image, 100% of english speakers responded with thumb — a specific term. in contrast, spanish speakers were not unanimous in their responses. figure 2 presents the production results for digits on the hand. while the single-word translational equivalent to ‘thumb’, pulgar, was preferred, only approximately 42% of participants used it. mano – spanish for ‘hand’ – was the second most frequent (17.4%). spanish speakers preferred using specific terms 63.2% of the time. for the ring finger images (presented in figure 2), english speakers preferred the specific term ring finger 83% of the time, but some participants (17%) did produce the general term finger. again, spanish speakers presented more variation in their responses than english participants, with seven unique utterances produced. 35% of spanish participants produced dedo anular (‘ring finger’); 26% produced dedo (‘finger’). overall, spanish participants trended like english participants with preference for specific (52.2%) over general terms (47.8%), although the preference was not as strong (see figure 2). proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 149 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2: english and spanish production of finger terms;“specific” terms refer to a single digit proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 150 https://doi.org/10.3765/elm https://www.elm-conference.net/ english production for the pinky was similar to english production for the ring finger (see figure 2): there is variation between specific (e.g. pinky – 45.8%, or pinky finger – 41.7%) and general term usage (12.5% finger). the single word pinky is used (and most frequently), but almost as many participants chose to use the two-word alternative pinky finger. spanish speakers also gave varying responses. the most common was the single-word equivalent for ‘pinky’, meñique (39.1%), vs. 21.7% for dedo ‘finger’. for the big toe images, english speakers favored the specific term big toe (83.3%) over the general term toe (16.7%). in contrast, spanish participants produced six different descriptions (shown in figure 3). the most common response (69.6%) was the specific term dedo gordo del pie (lit. ‘fat digit of the foot’), or ‘big toe’. the second most common response (34.8%) was the general term dedo del pie (lit. ‘digit of the foot’), which translates to toe. for the ring toe image, english participants showed increased dispersion in their responses. the majority of participants (58.3%) used the general term “toe”. the remaining participants (41.7%) produced various specific terms for the digit (“fourth toe” and “ring toe” to name a few). spanish speakers had a much higher rate of general term usage, with 82.6% of participants preferring terms like dedo del pie ‘toe’, dedo ‘digit’, or pie ‘foot’. finally, the production results for the pinky toe were similar between english and spanish. english speakers preferred using the general term toe far less than a specific term (20.8% and 79.2%, respectively). there was less dispersion in english production results for the pinky toe than for the ring toe. in spanish, participants generally preferred a specific term (52.2%) over a general term (47.8%), but the trend was not as strong as in english. in contrast to english, spanish production data exhibited a much larger amount of dispersion for the pinky toe than the ring toe, as shown in figure 3. these production results support spanish speakers’ intuitions that the spanish single-word alternative pulgar ‘thumb’ is less prevalent than thumb is in english. if geurts (2011) is right, then spanish speakers and english speakers should differ in scalar implicature calculation due to the differences in prevalence of the alternative forms for ‘thumb’. 3. comprehension studies. we are now in a position to address our main research question: upon hearing a general term for a digit (e.g. finger or toe), what alternatives do english and spanish speakers use to compute alternatives? are alternatives activated in accordance with their salience (as measured by production frequency in our production experiments) or in accordance with their complexity (as measured by number of words)? although these two things are correlated, they are not identical. for example, if geurts (2011) is right, and the salience of alternatives matters, then english and spanish will differ with respect to the scalar implicature associated with finger due to the difference in production probability for the more specific forms (thumb vs. pulgar). if complexity is all that matters, then there should be an implicature in both languages, because there is an equally simpler, yet more informative alternative in both languages. through our comprehension studies, we are able to distinguish among these hypotheses. 3.1. participants. 45 american english participants and 48 mexican spanish participants, recruited via prolific, completed the comprehension task. english participants were self-reporting american monolinguals that were born and currently reside in the united states. spanish participants were self-reporting mexican monolinguals that were born and currently reside in mexico. proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 151 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 3: english and spanish production of toe terms; “specific” terms refer to a single digit proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 152 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.2. materials. the target items for the comprehension studies consisted of 6 image pairs. three of the pairs were images of hands and three of the pairs were images of feet such that all possible hand combinations and all possible foot combinations were presented. no target pairs consisted of an image of a digit on the hand and an image of a digit on the foot. the images were the same six images from the production study (see figure 1). in addition to the 6 target image pairs, participants were also presented with 6 filler image pairs. three of the filler pairs were “easy”, where the utterance clearly matched only one of the images (e.g. she has a tattoo on her back, with a pair of images that contained only one back tattoo). the other three filler pairs were considered “hard”; these image pairs contained, for example, two different back tattoos. filler pairs that were “easy” acted as attention checks, since there was a clear correct response. participants who failed one or more “easy” fillers were eliminated from the results. 3.3. procedure. on each trial, a pair of images was presented, both showing a tattoo on a body part. on critical trials, the images showed tattoos on two different fingers, or two different toes: thumb on the left, ring finger on the right, for example. along with the images, participants read an utterance of the form she has a tattoo on her x, where x was a general term: finger or toe or the spanish translational equivalent (dedo or dedo del pie). participants were asked “which picture are they talking about?” and clicked on an image. item order and left-right presentation of the images were randomized. 3.4. results. responses were simply coded as the image the participant clicked on (e.g. “thumb” for the image with the tattoo on the thumb). the p-values we report are the result of conducting a 1-sample proportion test, where the null hypothesis, or the probability of choosing the correct image is 0.5. the assumption is that the data follow a bernoulli distribution. we ran a benjaminihochberg adjustment on the p-values. in the comprehension study, when participants were asked to choose between the thumb image and the ring finger image given the statement “she has a tattoo on her finger”, 75% of english participants chose the image of the ring finger, p = 0.004 (see figure 4). in contrast, just over half of the spanish speakers chose the ring finger image over the thumb image, but the error bar, which depicts a 95% confidence interval, distinctly crosses the 50% mark, showing that the spanish participants’ responses are not statistically significantly different from chance (p = 0.627). for the big toe and ring toe image pair, english participants showed a slight preference for the ring toe image, with roughly 63% of participants choosing that image. however, the error bar indicates that this result is not statistically different from chance (p = 0.145). spanish participants actually showed a stronger trend toward the ring toe given the translational equivalent of “she has a tattoo on her toe” (p = 0.007). a full summary of the results is given in table 1. the estimate is the estimated true proportion in the greater population. we conducted a benjamini-hochberg adjustment to obtain the adjusted p-values. what we can take away from these results is that the following implicatures exist in english: • finger ↝ not thumb • toe ↝ not pinky toe proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 153 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 4: observed frequency with 95% ci of choosing thumb or ring finger in english and spanish. in spanish, we have: • dedo del pie ↝ not pinky toe • dedo del pie ↝ not ring toe not all of these implicatures are predicted by the assumption that complexity alone is what drives the activation of alternatives. in the next section, we develop precise computational models in order to understand the significance of these results more deeply. 4. bayesian modeling. 4.1. model definitions. to gain a better understanding of complexity and prevalence, and their roles in scalar implicature, we compared two rational speech act models (frank & goodman 2012, goodman & stuhlmüller 2013; i.a.) that differ in how the speaker is defined. the first model incorporates a traditional speaker model that penalizes longer – more complex – utterances (henceforth referred to as the complexity model). the second model is a prevalence-based speaker model that has perfect knowledge of speaker production (henceforth referred to as the production model). for both models, the space of possible states includes six underlying states, corresponding to the six target digits (thumb, ring finger, pinky, big toe, ring toe, and pinky toe). literal meanings for each utterance from the production study were hand-specified as a subset of the states. for the complexity-based speaker model, as presented earlier, the speaker chooses an utterance based on accuracy and cost. length is equivalent to length in words, and l0(s ∣u) is the probability that a literal listener will choose a state s given an utterance u. the model contains proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 154 https://doi.org/10.3765/elm https://www.elm-conference.net/ condition language estimate p-value adj. p-value 1 big toe vs. ring toe eng 0.64 0.097 0.145 2 big toe vs. ring toe spa 0.72 0.002 0.007* 3 big toe vs. pinky toe eng 0.34 0.05 0.10 4 big toe vs. pinky toe spa 0.50 1.00 1.00 5 ring toe vs. pinky toe eng 0.24 0.0006 0.002* 6 ring toe vs. pinky toe spa 0.21 0.00009 0.0009* 7 thumb vs. ring finger eng 0.75 0.002 0.004* 8 thumb vs. ring finger spa 0.56 0.47 0.627 9 thumb vs. pinky finger eng 0.80 0.0001 0.0009* 10 thumb vs. pinky finger spa 0.63 0.086 0.145 11 ring finger vs. pinky finger eng 0.45 0.651 0.781 12 ring finger vs. pinky finger spa 0.52 0.885 0.965 table 1: p-values and adjusted p-values for each language/condition pair. two free parameters. alpha (α) is the ‘rationality parameter’, which corresponds to how much the speaker maximizes utility, where utility in this context corresponds to accuracy, that is, probability that the literal listener selects the correct referent. the parameter β is a multiplier on cost, where cost is measured as number of words in the utterance. the cost parameter reflects speakers’ degree of preference to be as concise as possible when speaking. model parameters for the complexity model were tuned to the thumb/ring finger data point from the experimental results of the english comprehension study: α set at 1 and β set at 2. s(u ∣ s)∝ exp(α ⋅l0(s ∣u) − β ⋅ length(u)) a pragmatic listener was then built on top of the complexity speaker model. as usual in rsa, a pragmatic listener chooses an interpretation using bayes’ rule, reasoning about the likelihood that a speaker would choose various utterances under various hypotheses about what the speaker intends. l(s ∣u)∝ s(u ∣ s) ⋅ p (s) in contrast to the complexity model, the production model is fed the exact production probabilities for each utterance collected from the production study. the speaker in this model chooses an utterance based on the empirically observed frequencies in my production data. we write f (u ∣ s) to denote the frequency with which an utterance u was used in the production experiments to describe state s (i.e. the finger or toe that had the tattoo). s(u ∣ s)∝ f (u ∣ s) this ensures that the production model has full awareness of what utterances are more or less prevalent for speakers – these are utterances speakers actually produced. as in the complexity model, a pragmatic listener is coded on top of the production model. in fact, the pragmatic listener model is the same for both speaker models. because of the difference in the way that the speaker proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 155 https://doi.org/10.3765/elm https://www.elm-conference.net/ complexity model production model figure 5: model predictions plotted against comprehension results; inaccurate model predictions are circled in red. is defined, however, the pragmatic listener in the production model has full awareness of actual production probabilities. 4.2. model performance. the model predictions are presented alongside the empirical results in figure 5. the complexity model inaccurately predicts no implicature for the thumb/pinky item in english. this is because the model is only considering the fact that these two utterances are equally complex. additionally, the model incorrectly predicts that there will be an implicature for ring finger/pinky in english since the one-word term pinky is an available alternative. for digits on the feet, the complexity model incorrectly predicts no implicature for ring toe/pinky toe in english and big toe/ring toe in spanish – since they are equally complex. however, the production results suggest that they are not equally viable as alternatives. since ring toe and pinky toe are equally as complex, if complexity alone determined which alternatives were available to speakers, we would expect no implicature to arise. however, the presence of the implicature toe proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 156 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 6: comparison of model results ↝ ‘not pinky toe’ suggests that pinky toe is a more prevalent alternative than ring toe. these results suggest that something else is going on in calculating implicatures that complexity alone cannot account for. the complexity model fails to understand that two alternatives with equal complexity may differ in their prevalence. the production model, on the other hand, is capable of accounting for this. overall, the production model performs much better than the complexity model; see figure 5. the only incorrect prediction it makes is that there would be no implicature for big toe/pinky toe in english. the production model predicts stronger implicatures for the thumb/ring finger items in english and spanish, and for the thumb/pinky finger items in english. in other words, where the comprehension results trend rightward, at just over 75%, for the thumb/pinky item in english, the production model predicts 100% of participants selecting the pinky finger over the thumb. otherwise, the model predictions fall in line with all empirical results for the comprehension experiments in spanish and english. figure 6 plots the rate at which listeners chose the image on the right along the x-axis against the probability assigned to the image on the right by each model on the y-axis. a perfect model would assign probability at the exact same rate as actual production. the r2 for the complexity model is only 30.6%. this means that the complexity model accounts for 30.6% of the variation present in the comprehension studies. the r2 for the production model, in contrast, is 76.3%, which is to say that the production model accounts for 76.3% of the variance in the data. there is a stark contrast in the explanatory power of each model. this shows that listeners have a good mental model of speakers, and that their mental model is not purely complexity-based. in fact, the comparison of these model results suggests that speakers are considering prevalence over complexity, since the prevalence-based speaker model does not include a cost parameter. proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 157 https://doi.org/10.3765/elm https://www.elm-conference.net/ 5. conclusions. the results outlined above suggest that spanish and english speakers do differ with respect to the scalar implicatures associated with finger in accordance with the prevalence of the words for ‘thumb’, ‘ring finger’ and ‘pinky finger’. our empirical results support the idea that differences across languages in the implicatures associated with general terms are closely tied to differences in production probabilities for more specific terms. since pulgar ‘thumb’ is not as prevalent in spanish as thumb is in english, it is not available in the set of alternatives to finger, which is why speakers do not calculate an implicature. our modelling results further support the conclusion that alternatives are constrained based on prevalence: the production model significantly outperforms the complexity model. while complexity does assist in determining the set of alternatives present for speakers, it is not as explanatory as full awareness of what speakers actually produce. prima facie, these findings go against structural theories, like katzir (2007) and horn (2000), that constrain the set of alternatives based on complexity alone. these results also support the idea that activation of alternatives is not an all-or-nothing matter, and this is an idea that is naturally captured in a bayesian pragmatic model that relies on gradient speaker production probabilities. references frank, michael c. & noah d. goodman. 2012. predicting pragmatic reasoning in language games. science 336(6084). 998. geurts, bart. 2011. quantity implicatures. cambridge: cambridge university press. goodman, noah d. & andreas stuhlmüller. 2013. knowledge and implicature: modeling language understanding as social cognition. topics in cognitive science 5(1). 173–184. horn, lawrence r. 2000. from if to iff: conditional perfection as pragmatic strengthening. journal of pragmatics 32(3). 289–326. katzir, roni. 2007. structurally-defined alternatives. linguistics and philosophy 30(6). 669–690. proceedings of elm 1: 147-158, 2021 danielle dionne and elizabeth coppock: tattoos as a window onto cross-linguistic differences in scalar implicature. 158 https://doi.org/10.3765/elm https://www.elm-conference.net/ semantics of non-doxastic attitude ascriptions from experimental perspective wojciech rostworowski, katarzyna kuś & bartosz maćkiewicz* abstract: the paper presents novel experimental data regarding reports of nondoxastic attitudes (expressed by verbs such as “wants”, “fear”, “is glad”, and etc.) as observed by some theorists, non-doxastic attitude ascriptions differ from the ascriptions of doxastic attitudes (e.g., “believes”) in that they do not support simple entailments or presuppositions of their complement clause. in particular, an ascription may intuitively change its truth-value if we alter the informational structure of the embedded clause without modifying its truth conditions. we present two experiments whose results support this observation. experiment 1 shows that the truth-value and acceptability judgements of non-doxastic attitude ascriptions in a context generally depend on the informational structure of the embedded clause. experiment 2 reveals that the truth-value judgements vary if we manipulate not only the “presuppositionassertion” structure of the embedded clause, but also the components related to the non-presuppositional entailments of the clause. this conclusion suggests that the contents on which attitude verbs operate should be represented as structured entities. keywords. attitude ascriptions; entailments; non-doxastic attitudes; presuppositions; semantics; truth-value judgements 1. introduction. a number of theorists (e.g., heim 1992, elbourne 2010, maier 2015, rostworowski 2018) have observed that there is a certain asymmetry between reports of doxastic attitudes – like beliefs or knowledge – and reports of the non-doxastic attitudes – like desires, fears, feeling glad, etc. consider the following pairs of ascriptions: (1) a. anne believes that the ghost from the attic is quiet. ⇒ b. anne believes that there is a (unique) ghost in the attic and it is quiet. (2) a. anne is glad that the ghost from the attic is quiet. ⇏ b. anne is glad that there is a (unique) ghost in the attic and it is quiet. (e.g., elbourne 2010) provided we regard (1a) as being de dicto (i.e., we take “the ghost” to be inside the scope of the belief-operator at the level of the sentence logical form), the ascription has the reading which trivially entails (1b). in general, two ascriptions in (1) seem to be equivalent given that their embedded clauses have closely related contents. on the other hand, (2b) is essentially different from (2a) and can be intuitively false in a situation in which (2a) is intuitively true (e.g., when anne is glad that the ghost is quiet but not happy that the ghost exists at all). so, (2a) and (2b) have different truth conditions. yet, the difference between the complement clauses in (2) is exactly the same as in (1). we will use the term “hyperintensionality” to refer to the indicated feature of non-doxastic attitude ascriptions. the paper aims to contribute experimental data regarding the hyperintensionality of nondoxastic attitude ascriptions. first, we seek to verify whether ordinary users of language share the * the research is supported by the grant of the national science center in poland no. 2020/37/b/hs1/01605. we are grateful to the audience of experiments in linguistic meaning 2 for their feedback about our project. authors: wojciech rostworowski, university of warsaw (w.rostworowski@uw.edu.pl), katarzyna kuś, university of warsaw (kkus@uw.edu.pl) & bartosz maćkiewicz, university of warsaw (b.mackiewicz@uw.edu.pl). proceedings of elm 2: 241-251, 2023 c©2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz published by the lsa with permission of the author(s) under a cc by license. 241 https://doi.org/10.3765/elm https://www.elm-conference.net/ intuition that ascriptions of non-doxastic attitudes may have different truth values, even though their complement clauses have the same truth-conditional contents. second, we want to investigate what kinds of manipulations of the complement clause structure are responsible for the expected difference in the truth-value evaluations. in section 2, we provide some theoretical considerations about what is potentially responsible for the asymmetry in the truth-value evaluations of ascriptions such as (2a) vs. (2b). in section 3, we present our experiments. section 4 offers a discussion of our results and indicates further research directions. 2. attitude verbs, presuppositions, and entailments. let us take a closer look at the problem of hyperintensionality. assuming that attitude verbs operate on the contents of their embedded clauses and nothing else – in line with compositional semantics – the contrast between (2a) and (2b) must be regarded as the output of a difference between the contents of the embedded clauses. one possible explanation of the contrast is that the embedded clauses – i.e., “the ghost from the attic is quiet” and “there is a (unique) ghost in the attic and it is quiet” – have different informational structures in the sense that the first one presupposes the existence of a ghost, while the second explicitly asserts it. in other terms, there is a difference between the sets of presuppositions of the sentences that serve as the complement clauses in (2a) and (2b).1 there are other examples which confirm the prediction that incorporating the presupposition of the embedded clause as a part of its assertoric content affects the intuitive interpretation of the whole sentence. consider: (3) a. john wishes that anne would quit smoking. ⇏ b. john wishes that anne would smoke and quit it. (4) a. jane wonders whether it was jones who murdered smith. ⇏ b. jane wonders whether smith was murdered and it was jones who had done it. on the most natural readings, (a)-ascriptions express different contents from (b)-ascriptions. the appeal to presuppositions can explain the effect of “hyperintensionality” related to nondoxastic attitude ascriptions. a prominent feature of presuppositions is that they project in various kinds of embeddings, i.e., they continue to arise if the trigger is embedded under certain operators (for example, negation, or the antecedent of a conditional; see karttunen 1974). when projecting, the presupposition escapes the scope of a given operator at the same time. this phenomenon can be illustrated with a simpler example including negation: (5) it wasn’t jones who murdered smith. a natural reading of the sentence is the one on which it (still) presupposes the existence of smith’s murderer and denies that jones is the murderer. that is, what is denied is only the fact that jones murdered smith – and not the fact that smith has been murdered at all. the projection behavior of presuppositions whose triggers are embedded under attitude verbs is more complex. yet, different analyses of this behavior agree with some general observations and predictions, which roughly fit the pattern illustrated by (5). namely, the ascription such as (2a) – which contains a presuppositional trigger in the complement clause (here, ‘the ghost from the attic’) – presupposes as a whole that either the local presupposition is simply satisfied, or that the attitude holder believes 1 for this reason, some theorists (e.g., elbourne 2010, 2013) have argued that the contrast between (2a) and (2b) and the like provides an argument for the presuppositional treatment of definite descriptions and against russellian analysis on which descriptions are treated as quantifiers. arguably, russell’s theory is committed to the claim that (2a) (on the de dicto interpretation) has the reading exactly equivalent to (2b). (for criticisms of this claim, see kaplan 2005, neale 2005, pupa 2013, rostworowski 2018). proceedings of elm 2: 241-251, 2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz: semantics of non-doxastic attitude ascriptions from experimental perspective. 242 https://doi.org/10.3765/elm https://www.elm-conference.net/ that it is satisfied (see karttunen 1974, heim 1992, geurts 1998, maier 2015). in particular, (2a) presupposes that anne believes that the ghost from the attic exists. the mere presupposition of the ghost’s existence is at the same time analyzed as escaping the scope of the non-doxastic attitude verb.2 so, (2a) does not state, among other things, that anne is glad that the ghost exists. on the other hand, when the existence of the ghost is a part of the asserted content – like in (2b) – it naturally falls under the scope of the non-doxastic attitude verb on the de dicto interpretation. consequently, (2b) expresses a different attitude than (2a). in short, the presuppositional account can predict that (2a) and (2b) are non-equivalent. however, some further examples of non-doxastic attitude ascriptions (see blumberg 2017, rostworowski 2018) provide evidence for the hypothesis that hyperintensionality does not manifest itself only in the cases of presuppositional differences between the complement clauses in the ascriptions. consider the following examples: (6) a. jessica wants to buy the one-thousand-dollar necklace. ⇏? b. jessica wants to spend one thousand dollars and buy the necklace for this amount. (7) a. brad wonders whether the dictator has been assassinated. ⇏? b. brad wonders whether the dictator is dead and has been assassinated. (rostworowski 2018) arguably, (a)-ascriptions express somewhat different attitudes than (b)-ascriptions; in particular, we may imagine a context in which (a)-ascriptions are intuitively correct and (b)-ascriptions are not. at the same time we can observe that the first conjunct of the complement clause in the above (b)-ascriptions should not be regarded as a presupposition of the corresponding complement clause in (a)-ascription – just like it is in the earlier examples (2)-(4). for instance, the statement of “jessica bought the one-thousand dollar necklace” entails that jessica spent one thousand dollars (and bought the necklace for this amount), but the latter is not presupposed by the former in the technical sense. in particular, this sort of entailment does not exhibit the proper projection behavior, which, according to the earlier consideration, is a typical feature of presuppositions. to sum up, non-doxastic attitude ascriptions seem to display hyperintensionality if the complement clauses induce structural differences of the non-presuppositional nature, likewise in the case in which the embedded clauses have different presuppositions. that is to say, the intuitive truth-value of a non-doxastic attitude ascription depends on the way how the information in the complement clause is structured, in addition to its truth condition. hence, two ascriptions where the complement clauses are truth-conditionally equivalent may be evaluated differently in a context. finally, hyperintensionality of non-doxastic attitude ascriptions provides a challenge for semantic theory. roughly speaking, it shows that the semantic content of a sentence – which serves as the input for attitude-verbs operators – is fine-grained to the extent that the sentence informational structure must be included in the content representation. in light of this, the nonstructural notions of content – like the ones defining it in terms of sets of possible worlds or situations – require substantive revisions in order to handle non-doxastic attitude verbs properly. in particular, it might prove difficult to develop a theory which predicts that, for instance, (7a) is true while (7b) is not in the same context, provided that the state of the dictator’s being dead is intuitively a part of any situation in which the dictator has been assassinated. (for some discussions and proposals see, e.g., roelofsen & uegaki 2016, blumberg 2017). 2 for instance, this prediction is entailed by the analyses of geurts (1998) and maier (2015). proceedings of elm 2: 241-251, 2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz: semantics of non-doxastic attitude ascriptions from experimental perspective. 243 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3. experimental studies. the aim of our experimental studies was to provide empirical evidence for the theoretical claims made above. in experiment 1, we investigated whether the general informational-structure manipulations of the complement clause affected the truth-value or acceptability judgements of an attitude ascription in a context. experiment 2 further explored the issue by testing whether the expected difference in evaluations depends on the nature of the manipulation – on whether the manipulated element of the complement clause has a presuppositional nature or is a non-presuppositional entailment. 3.1. experiment 1. in this experiment, we compared two kinds of non-doxastic attitude ascriptions: straight ascriptions and complex ascriptions. the former embedded simple propositions as their complement clauses (cf. 2a, 3a), while the latter a conjunction of two claims (cf. 2b, 3b). in complex ascriptions, the attitude ascribed to a protagonist consisted of two parts: the first one explicitly stating a presuppositional or entailed content of the straight ascription, the second being the straight ascription itself. here are some examples used in our experiment: (8) a. linda is glad that her presentation today convinced the client. (straight ascription) b. linda is glad that she had a presentation today and convinced the client. (complex ascription) (9) a. linda fears that her presentation will not convince the client to sign the contract. (straight ascription) b. linda fears that she will have a presentation and she won’t convince the client to sign the contract. (complex ascription) due to the exploratory nature of the study, the non-doxastic attitudes we studied were diverse in nature. two of them were expressed by a pro-attitude verb (“want” and “glad”), two were conattitudes (“fear” and “feel sorry”). in addition, our verbs differed grammatically because “want”, a basic pro-attitude verb, does not introduce the content of the attitude with a that-clause but needs a verb in the to-infinitive form in the complement. the other verbs we have chosen are complemented by a that-clause, although the verb “glad” can be also followed by a verb in the toinfinitive form. in line with our theoretical considerations, we predicted differences between the evaluations of straight and complex ascriptions. in particular, people should tend to evaluate straight ascriptions more positively than complex ascriptions in the contexts presented in our questionnaires. 3.1.1. methods. we used questionnaires with an acceptability task and a truth-value judgment task. after reading a short fictional story, the respondents evaluated either a straight ascription or a complex one. they assessed either the truth value of a given ascription or its acceptability. both types of questions were followed by a confidence level question regarding the given answer. the participants answered on a 100-point visual analogue scale in the form of a slider ranging from 50 (“strongly disagree”) to +50 (“strongly agree”). in addition, they were asked a categorical comprehension question which appeared in advance to the critical evaluation question. the exact form of the questions is presented in table 1. proceedings of elm 2: 241-251, 2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz: semantics of non-doxastic attitude ascriptions from experimental perspective. 244 https://doi.org/10.3765/elm https://www.elm-conference.net/ questions truth-value judgment acceptability comprehension question “to whom linda had to give the presentation?” (“a client” and “an employee”) main question “in the light of the story, how would you evaluate the statement a?” (“true” and “false”) “in the light of the story, would you accept the statement a?” (“yes” and “no”) confidence level (if they answered true) “to what extent do you agree that the sentence a is true?” (if they answered fasle) “to what extent do you agree that the sentence a is false?” (if they answered yes) “to what extent do you agree that the sentence a is acceptable?” (if they answered no) “to what extent do you agree that the sentence a is unacceptable?” table 1: example prompts (a stands for an attitude ascription) to sum up, we employed a 2 x 2 x 4 mixed experimental design with two between-subject factors: a type of ascription (straight vs. complex) and a type of a task (acceptability vs. truth-value judgment), and one within-subject factor (verb: “fear” vs. “want” vs. “feel sorry” vs. “glad”). 3.1.2. participants. for experiment 1, 331 participants were recruited on clickworker to complete an online questionnaire. non-native english speakers and subjects who failed the attention or comprehension check were excluded. the final sample consisted of 285 subjects (163 females, 120 males, one person who refused to answer and one person who chose “other”; mean age: 39.35). 3.1.3. materials. we developed four kinds of situations that differ in their setup and plots, and named them after the protagonists: linda, anne, mark, and john. the difference across setups was not predicted to be significant and served merely as a robustness check. each vignette depicted a situation in which non-doxastic attitudes of the protagonist can be accurately/truly described by a straight ascription. however, certain beliefs and other attitudes of the protagonist indicated that the first conjunct of the embedded conjunction in the complex ascription was not a part of the protagonist’s attitude. table 2 presents a sample of the mark vignette in all four conditions (four non-doxastic verbs) with both straight and complex ascriptions evaluated by the study participants. 3.1.4. procedure. the study had the form of an on-line questionnaire in which participants were asked to read a total of eight short fictional stories (“vignettes”, 4 targets + 4 fillers) and answer questions about them. each participant was randomly assigned to two between-subject conditions (questions on a straight/complex ascription and the truth-value judgment/acceptability rating task) and received four vignettes with different setups, one for each within-subject condition (nondoxastic attitude verbs). four setups (linda, anne, mark, john) were counterbalanced across presentation lists in such a way that no participant received two vignettes with the same setup while each setup was equally likely to occur with any non-doxastic attitude verb. additionally, four experimental vignettes were interspersed with four vignette-fillers. proceedings of elm 2: 241-251, 2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz: semantics of non-doxastic attitude ascriptions from experimental perspective. 245 https://doi.org/10.3765/elm https://www.elm-conference.net/ want fear feel sorry glad mark is a third-year student of history. he attends a philosophy course which he doesn’t enjoy very much. he prefers online meetings over the regular stationary ones. mark wishes the continuation of the philosophy course next year would be online. mark is a third-year student of history. he attends a philosophy course which he enjoys very much and wants to continue next year. but he prefers regular stationary meetings over the current online ones. as the pandemic grows, marks is afraid that next year the classes will be continued only online. mark is a third-year student of history. he attends a philosophy course which he enjoys very much and wants to continue next year. but he does not like the online form of studying and so is not happy that the philosophy classes are conducted on the internet this year. mark is a third-year student of history. he attends a philosophy course which he doesn’t enjoy very much. he prefers online studying over the regular stationary meetings, especially when it comes to the philosophy course. this year the university runs all courses online due to the pandemic, which makes mark happy. (straight ascription) mark wants the philosophy course next year to be online. (straight ascription) mark fears that the philosophy classes next year will be online. (straight ascription) mark feels sorry that the philosophy classes he attended this year are online. (straight ascription) mark is glad that the philosophy classes this year are online. (complex ascription) mark wants a philosophy course next year and for it to be online. (complex ascription) mark fears that there will be philosophy classes next year and they will be online. (complex ascription) mark feels sorry that he has attended philosophy classes this year and they have been online. (complex ascription) mark is glad to have philosophy classes this year and that they are online. table 2: example stimuli (experiment 1) 3.1.5. results. in order to analyze the data, we computed a compound index for each rated ascription in the following way: we took confidence rating and multiplied it by 1 if the answer to the main question was positive (true or yes) and by -1 if it was negative (false or no). using a 2 x 2 x 4 anova, we found a statistically significant effect of the ascription type (f(1, 281) = 116.3, p < 0.001, η2 g = 0.131). in line with our predictions, straight ascriptions (acceptability rating: m = 35.1, sd = 12.3; truth-value judgment: m = 35.9, sd = 12.8) were rated substantially higher compared to the complex ones (acceptability rating: m = 16.3, sd = 19.6; truth-value judgment: m = 14.2, sd = 17.0). figure 1 shows the mean evaluations of ascriptions in each condition. figure 2 shows the means for all attitude verbs jointly. no statistically significant difference between the two types of task was observed (f(1, 281) = 0.170, p = 0.68). the only statistically significant interaction was between the ascription type and verb factors (f(3, 834) = 7.147, p < 0.001, η2 g = 0.016). a closer examination of the data revealed that for each verb, the difference between straight and complex ascriptions was statistically significant (p < 0.01), but for “fear” and “glad” the difference was less pronounced compared to “glad” and “feel sorry”. proceedings of elm 2: 241-251, 2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz: semantics of non-doxastic attitude ascriptions from experimental perspective. 246 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: experiment 1 means by condition figure 2: experiment 1 means by condition for all verbs collapsed 3.2. experiment 2. the second experiment aimed to verify whether the difference in evaluations between straight and complex ascriptions persists both in the case in which the clause embedded in the complex ascription explicitly formulates its presupposition, as well as a nonpresuppositional entailment. we hypothesized that straight ascriptions are evaluated more positively than the complex ones in both cases. 3.2.1. methods. due to the lack of differences in the type of task in the first experiment, we used only the truth-value judgment tasks. this is the only change in methods compared to experiment 1. we employed a 2 x 2 x 2 mixed experimental design with one between-subject factor, i.e., a type of ascription (straight vs. complex) and two within-subject factors: a type of informationstructure manipulation (presupposition vs. entailment) and attitude verb (“want” vs. “glad”). 3.2.2. participants. in total, 352 participants recruited on clickworker took part in the study. after excluding those who failed the attention check at the beginning of the survey, 292 participants remained. then those who failed at least one comprehension check concerning target stories were also excluded. 290 subjects remained in our sample (196 females, 89 males, five persons who chose “other”; mean age: 36.23). 3.2.3. materials. again, we developed four kinds of setups that differ in their plots and were named after the protagonists: andrew, tanja, arthur and jessica. the idea behind the stories was the same as in experiment 1: while the protagonist’s attitude could be accurately described with a straight ascription, it was questionable whether it could be described with a corresponding complex ascription. for each setup, we proposed two types of complex ascriptions: the first conjunct was either a presupposition, or a non-presupposed entailment of the second conjunct (i.e., the presupposition and entailment conditions). table 3 presents a sample of the andrew vignette in all four conditions with both straight and complex ascriptions evaluated by the study participants. 3.2.4. procedure. again, the participants were presented with an on-line questionnaire and asked to read a total of four target vignettes interspersed with four filler items. each participant was randomly assigned to one of two between-subject conditions (straight or complex ascription) and received four vignettes with different setups, one for each within-subject condition ( “want” vs. “glad”) and one for each type of the information-structure manipulation (presupposition vs. entailment). four setups were counterbalanced across presentation lists. proceedings of elm 2: 241-251, 2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz: semantics of non-doxastic attitude ascriptions from experimental perspective. 247 https://doi.org/10.3765/elm https://www.elm-conference.net/ want glad presupposition entailment presupposition entailment andrew owns a big old house. unfortunately, it has been seriously damaged by a huge flood. as costs of keeping the house have started to go up, andrew would rather not have to deal with this or any other house. he believes it would be better for him to sell the house and move to a smaller apartment. but before putting his house on the market, andrew decides to make some repairs and spruce his home up a little. andrew owns an old beautiful cottage. unfortunately, he will have to sell it because of high costs of running and keeping it in good condition. andrew generally does not want to sell the house because he really likes it. if he is to sell it to anyone, it would be to his cousin jennifer, an interior designer, since he believes that she will take good care of the cottage. andrew owns a big old house. unfortunately, it has been seriously damaged by a huge flood. andrew decides to make some repairs. however, afterwards, he feels uneasy living there because he is scared that something like that would happen again. he then decides to sell this house and move to a small apartment that will not be as troublesome as a house. when putting his house on the market, he realizes that the improvements he has made after the flood have significantly increased the value of his property, which makes him really happy. andrew owned an old beautiful cottage. unfortunately, he had to sell it because of high costs of running and keeping it in good condition. andrew is generally not happy about selling the house because he really liked it. however, at the same time he feels pleased that he has sold it to his cousin jennifer, an interior designer, since he knows that she will take good care of the cottage. (straight ascription) andrew wants to renovate his house. (straight ascription) andrew wants to sell his house to jennifer. (straight ascription) andrew is glad that he has renovated his house. (straight ascription) andrew is glad that he has sold his house to jennifer. (complex ascription) andrew wants to own a house and to renovate it. (complex ascription) andrew wants to sell his house and sell it to jennifer. (complex ascription) andrew is glad that he owned a house and renovated it. (complex ascription) andrew is glad that he has sold his house and that he has sold it to jennifer. table 3: example stimuli (experiment 2) 3.2.5. results. the preparation of the data was the same as in experiment 1. the results are presented in figure 3 which shows means by condition in experiment 2. the main finding of the first study was replicated. the participants rated straight ascriptions (presupposition: m = 35.94, sd = 16.00; entailment: m = 22.73, sd = 20.80) significantly higher than complex ones (presupposition: m = -12.60, sd = 26.50; entailment: m = -0.35, sd = 25.10; f(1, 576) = 373.1, p < 0.001, η2 g = 0.26). we also found a statistically significant effect of the verb (f(1, 576) = 78.84, p < 0.001, η2 g = 0.06). two interactions were statistically significant. the first one was a two-way interaction between the type of information-structure manipulation and type of ascription (f(1, proceedings of elm 2: 241-251, 2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz: semantics of non-doxastic attitude ascriptions from experimental perspective. 248 https://doi.org/10.3765/elm https://www.elm-conference.net/ 576) = 47.15, p < 0.001, η2 g = 0.04). the effect of the ascription type was more pronounced in the presupposition condition compared to the entailment. as can be seen in figure 4, the difference between straight and complex ascriptions is larger in the presupposition condition. the second significant interaction was a three-way interaction between the information-structure manipulation, a type of ascription and a verb (f(1, 576) = 7.384, p = 0.007; η2 g = 0.006). more detailed inspection of the data shows that the interaction effect is present due to the fact that in the presupposition condition, the difference between straight and complex ascriptions for “want” is larger than for “glad”. figure 3. experiment 2 means by condition figure 4. experiment 2 means by condition for all verbs collapsed 4. discussion and conclusions. the results of both experiments have confirmed our theoretical predictions. experiment 1 has shown that evaluations of the non-doxastic attitude ascriptions generally depend on the informational structure of the embedded clause. in particular, people tend to agree that straight ascriptions are “true” or accept them to a much greater degree than complex ascriptions in the contexts presented in our study. experiment 2 has provided a more specific insight into the phenomenon at issue. a subject s may hold an attitude towards q and not towards p (where p is a presupposition of q) and, in such a case, people tend to evaluate the ascription saying that s holds the given attitude towards the conjunction of p and q as false rather than true. it shows that presuppositions indeed escape the scope of attitude-verbs operators in the sentences like straight ascriptions. this observation is in line with predictions of the theoretical analyses of presupposition projection. however, experiment 2 also has demonstrated that a similar effect (though somewhat weaker) arises with non-presuppositional entailments. that is to say, according to the study participants, s may hold an attitude towards q but not towards the conjunction of p and q, even if p is entailed by q, so q and “p and q” are genuinely equivalent. this result indicates that non-doxastic attitude ascriptions are indeed sensitive to the way in which the content of the complement clause is structured, which means that non-doxastic attitude verbs operate on structured contents (rather than, e.g., sets of possible worlds). as we have observed, the difference between evaluations of straight and complex ascriptions was more pronounced in the presupposition condition than in the entailment one. in particular, people tended to evaluate complex ascriptions as closer to “false” rather than “true” in the presupposition condition, while they expressed genuinely ambivalent judgments in the entailment condition (rate of acceptance close to 0). our hypothesis is that this difference is due to pragmatics. proceedings of elm 2: 241-251, 2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz: semantics of non-doxastic attitude ascriptions from experimental perspective. 249 https://doi.org/10.3765/elm https://www.elm-conference.net/ an entailment being not presuppositional is practically something that follows from the assertoric part of a statement, thereby it is related to the statement main content in a more direct way than a presupposition. in a sense, such an entailment indicates a consequence of a given action, state, etc. expressed in the statement. presumably, when evaluating a sentence “s is glad that p and q”, or “s wants p and q”, some of the study participants felt that s had nonetheless accepted p as a consequence of the desirable q, although p itself was not desirable for s. consequently, these participants tended to agree that the complex ascriptions were true to a certain degree. to sum up, we have argued that non-doxastic attitude verbs are sensitive to the informational structure of the embedded clause. in the first part of the paper, we have presented theoretical observations in support of this claim. we have also argued that the phenomenon under discussion can be (only) partially explained by the theory of presuppositions and their projection. in the second part of the paper, we have provided novel experimental evidence that supports our theoretical observations. the results have shown that the informational structure of the embedded clause is relevant to the truth-value evaluation of a non-doxastic attitude ascription and confirmed the prediction that presuppositions project out of the scope of attitude verbs. there are further questions which arise with regards to our experimental findings. firstly, none of the conditions in which people evaluated complex ascriptions prompted them to give definitely negative answers. this raises the question of whether the rejection of complex ascriptions reflects a genuine semantic judgment, i.e., the ascriptions are indeed false according to people. if yes, then the respondents must have refused to express such a judgment in a definite way for some reasons (for instance, they observed that the second conjunct in the embedded clause correctly described the given attitude, so they wanted to deliver a less harsh verdict). the alternative is that the rejection is based on purely pragmatic grounds – that is, people regarded complex ascriptions as misleading but not literally false in the presented contexts. the second question is what kinds of entailments are actually supported by non-doxastic attitude verbs. the obtained results showing that straight ascriptions do not validate complex ascriptions indicate that at the same time conjunction elimination may work “under” an attitude verb (i.e., “s ves that p and q” ⇒ “s ves that p” is valid). the reason why people reject complex ascriptions is likely because the first conjunct in the embedded clause incorrectly describes an attitude of a subject. if so, people evaluate the whole ascription “s ves that p and q” based on their evaluation of “s ves that p”, which indicates that they implicitly perform conjunction elimination. both issues require further experimental research. references blumberg, kyle. 2017. ignorance implicatures and non-doxastic attitude verbs. proceedings of the 21st amsterdam colloquium. 135–144. https://semanticsarchive.net/archive/jzim2fhz/ac2017-proceedings.pdf elbourne, paul. 2010. the existence entailments of definite descriptions. linguistics and philosophy 33(1). 1–10. https://doi.org/10.1007/s10988-010-9072-3 elbourne, paul. 2013. definite descriptions. oxford: oxford university press. geurts, bart. 1998. presuppositions and anaphors in attitude contexts. linguistics and philosophy 21(6). 545–601. https://doi.org/10.1023/a:1005481821597 heim, irene. 1992. presupposition projection and the semantics of attitude verbs. journal of semantics 9(3). 183–221. https://doi.org/10.1093/jos/9.3.183 kaplan, david. 2005. reading “on denoting” on its centenary. mind 114. 933–1003. https://doi.org/10.1093/mind/fzi933 proceedings of elm 2: 241-251, 2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz: semantics of non-doxastic attitude ascriptions from experimental perspective. 250 https://doi.org/10.3765/elm https://www.elm-conference.net/ karttunen, lauri. 1974. presuppositions and linguistic context. theoretical linguistics 1. 181– 194. maier, emar. 2015. parasitic attitudes. linguistics and philosophy 38(3). 205–236. https://doi.org/10.1007/s10988-015-9174-z neale, stephen. 2005. a century later. mind 114. 809–871. https://doi.org/10.1093/mind/fzi809 pupa, francesco. 2013. embedded definite descriptions: a novel solution to the familiar problem. pacific philosophical quarterly 94(3). 290–314. https://doi.org/10.1111/papq.12001 roelofsen, floris & wataru uegaki. 2016. the distributive ignorance puzzle. proceedings of sinn und bedeutung 21. 999–1016. https://semanticsarchive.net/archive/gu1zte4z/paper.pdf rostworowski, wojciech. 2018. descriptions and non-doxastic attitude ascriptions. philosophical studies 175(6). 1311–1331. https://doi.org/10.1007/s11098-017-0912-7 proceedings of elm 2: 241-251, 2023 wojciech rostworowski, katarzyna kuś and bartosz maćkiewicz: semantics of non-doxastic attitude ascriptions from experimental perspective. 251 https://doi.org/10.3765/elm https://www.elm-conference.net/ overspecification of small cardinalities in reference production natalia zevakhina and elena pasalskaya* abstract. this paper presents experimental evidence for overspecification of small cardinalities in reference production. the idea is that when presented with a small set of unique objects (2, 3 or 4), the speaker includes a small cardinality while describing given objects, although it is overinformative for the hearer (e.g., “three stars”). on the contrary, when presented with a large set of unique objects, the speaker does not include cardinality in their description – so she produces a bare plural (e.g. “stars”). the effect of overspecifying small cardinalities resembles the effect of overspecifying color in reference production which has been extensively studied in recent years (cf. rubio-fernandez 2016, tarenskeen et al. 2015). when slides are flashed on the screen one by one, highlighted objects are still overspecified. we argue that one of the main reasons lies in subitizing effect, which is a human capacity to instantaneously grasp small cardinalities. keywords. overspecification; redundancy; informativeness; semantics; pragmatics; psycholinguistics; reference production; numerals; color adjectives 1. introduction. it is well-known that sometimes speakers tend to convey more information than required. in doing so, they use redundant linguistic expressions. imagine a situation where there is only one cup available in the visual context shared by the speaker and the addressee. in such a situation, the speaker utters (1), where she specifies the color of an object, even though the addressee can identify the object without paying attention to its color. in other words, the speaker could produce (2) where she only names the object itself without specifying its color. obviously, sentence (1) conveys more information than sentence (2). why did the speaker produce the overinformative sentence (1) instead of uttering the minimally informative sentence (2)? (1) give me the blue cup, please. (2) give me the cup, please. there have been proposed two accounts for this puzzle: speaker-oriented and hearer-oriented. according to the speaker-oriented account, producing overinformative utterances requires less cognitive effort from the speaker’s part than producing minimally informative utterances (pechmann 1989). to illustrate, in a situation of describing an object that possesses several attributes, the speaker need not think of which attributes are relevant to communicate. rather, she produces those attributes which have come to her mind and she does not waste time and efforts to make certain computations with respect to the relevance. according to the hearer-oriented account, * we would like to thank polina makarova, timofey morunov, veronika prigorkina, alina schipkova, elina sigdel for their help in searching for participants of both experiments and presenting them the materials. we acknowledge the discussion of hypotheses with alex dainiak and nina kazanina. we appreciate the remarks and suggestions from the audiences of the 8th experimental pragmatics conference (xprag) 2019 (university of edinburgh) and the audience of the 1st conference “experiments in linguistic meaning” (elm) 2020 (university of pennsylvania), online due to covid-19. authors: natalia zevakhina, hse university, russian federation (nzevakhina@hse.ru) & elena pasalskaya, hse university, russian federation (lena@ales.ru). this article is an output of a research project implemented as part of the basic research program at the national research university higher school of economics (hse university). proceedings of elm 1: 298-309, 2021 c©2021 natalia zevakhina and elena pasalskaya published by the lsa with permission of the author(s) under a cc by license. 298 https://doi.org/10.3765/elm https://www.elm-conference.net/ overspecification helps hearers identify relevant objects more rapidly, accurately and, therefore, more efficiently (mangold & pobel 1988, rubio-fernandez 2016). some attributes enable overspecification, whereas some others do not. recent research on attribute redundancy in reference production has demonstrated that color is much more likely to be used redundantly than size, material, shape, pattern, location and orientation (belke and meyer 2002, arts et al. 2011, brown-schmidt and konopka 2011, mangold and pobel 1988, gatt et al. 2007, gatt et al. 2013). the contrast between color and other attributes is accounted for in terms of absoluteness and salience (belke and meyer 2002, koolen et al. 2011, tarenskeen et al. 2015 among others). the absoluteness of color means that color is not contingent upon the speaker, the addressee, a situation, etc. for instance, if the observer sees a red object, she does not have to take into consideration colors of surrounding objects, she can merely report ‘red’. on the contrary, size is relative since the degree of size of a given object is evaluated by the observer among size degrees of surrounding objects. in case a given object does not differ from surrounding objects in terms of size, size is not likely to be reported. as for the salience, color is salient because of its high visual perceptibility. it is one of the features of preattentive analysis (trick 1992) and it is computed early in visual processing (livingstone and hubel 1988). other attributes (such as shape, material, pattern or size) are not so salient and, therefore, are less likely to be reported. moreover, the question which factor (absoluteness or salience) has a greater influence on the overspecification, seems to have a negative response: no factor. take size. it is relative. size is reported when it becomes contrastive and relevant for the communication (see van gompel et al. 2019). in other words, it is specified when it is supposed to be salient. for instance, the visual context includes only one big plate, with the rest of the plates being small. however, it is unlikely to be reported when the visual context includes three big and three small objects of one type or three big and three small objects of various types: a plate, a cup, a spoon, etc. moreover, the degree of the proportion between the objects does not seem to play a role here, as well. to illustrate, tarenskeen et al. (2015) manipulated size contrast between the objects with a proportion 3:1 but did not receive a significant increase in overspecification of size. now take pattern. it is absolute and salient. in this respect, it resembles color. however, gatt et al. (2013), tarensleen et al. (2015) showed that it is not as salient as color. to summarize, it seems that both factors – absoluteness and salience – determine overspecification in reference communication. focusing on color, it is noteworthy to mention several factors that affect color overspecification. firstly, color is more likely to be overspecified in polychrome contexts than in monochrome contexts (belke and meyer 2002, koolen et al. 2013, rubio-fernandez 2016). to illustrate, the speaker is more likely to say give me the blue cup, please in a situation when she is presented with objects of different colors than in a situation when she is presented with objects with the same color, say blue. secondly, color tends to be overspecified for atypically-colored objects in comparison to variably-colored or stereotypically-colored objects (westerbeek et al. 2015, rubio-fernandez 2016). to illustrate, the probability that the speaker redundantly utters (3) with an atypically-colored wolf is higher than the probability that the speaker redundantly utters (4) or (5) with a variably-colored car or a stereotypically-colored banana respectively. (3) show me a purple wolf, please. (4) show me a red car, please. (5) show me a yellow banana, please. proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 299 https://doi.org/10.3765/elm https://www.elm-conference.net/ thirdly, color is more often overspecified when it is variable than stereotypical for a given category of objects (sedivy 2003, rubio-fernandez 2016). for example, the speaker would produce (4) to a higher degree than (5). fourthly, color is more likely to be overspecified when referring to objects for which color is more important: e.g., artifacts like clothes or cars are more colorpertinent than geometric objects (rubio-fernandez 2016). going back to the distinction between color and size, brown-schmidt and konopka (2011) argued for that not only color but also number (or cardinality) is distinct from size. color adjectives and numerals were reported more frequently and faster than size adjectives in reference communication. they were reported both in contrastive and non-contrastive contexts. on the contrary, size adjectives were reported significantly more often in contrastive contexts. brownschmidt and konopka (2011) suggested that the reason for this is that, unlike color and number, size is a context-dependent modifier. these findings accord with the following idea circulated in the literature: size is relative (context-dependent), whereas color is absolute (not contextdependent); size is non-salient, whereas color is salient. as for the number, according to the findings by brown-schmidt and konopka (2011), it seems to be absolute and salient, like color. however, this fact was not directly addressed in the literature. furthermore, it was not clear why number seems to be absolute and salient. also, it was not clear which numbers were tested in brown-schmidt and konopka (2011). judging by figure (2a-b) in brown-schmidt and konopka (2011: 308), the numbers till 5 were involved. on the whole, it seems that what is needed now is a more systematic study of number overspecification in reference production. this is exactly what we do in the present study. 1.1. subitizing. for more than a century (since bourdon 1908), there has been acknowledged a fast, accurate and confident apprehension of cardinalities of small sets (1-4 or even 1-8) which considerably decreases in cardinalities of large sets. to illustrate, when presented with 3 dots in a display, an observer undoubtedly and very rapidly determines the cardinality of the dots. this capability vanishes when the observer is presented with 15 dots in a display. the phenomenon of immediately grasping the cardinality of few elements in a given set was coined as subitizing in kaufman et al. (1949), with a reference to dr. cornelina c. coulter (kaufman et al. 1949: 520) who suggested this term roughly meaning ‘sudden apprehension’, cf. also a similar term numerousness in stevens (1938), thomas et al. (1999) and taves (1941). in kaufman et al. (1949), subitizing was contrasted to estimation, that is, an approximate and less accurate apprehension of cardinalities of large sets, cf. also a close term numerosity in stevens (1938), thomas et al. (1999) and taves (1941). the question of what is the threshold for small sets is still debated but there has been a tacit agreement in the literature that the numbers from 1 to 3-4 belong to the subitizing range. whether the numbers from 5 to 8 belong to the subitizing range is still questionable. the threshold seems to vary from person to person. it is also dependent on a particular experiment setting and some other factors (akin and chase 1978, atkinson, campbell and francis 1976, chi and klahr 1975, jensen, reese and reese 1950, mandler and shebo 1982, oyama, kikuchi and ichihara 1981). importantly, subitizing is not the same as counting small cardinalities. rather, it involves a separate cognitive mechanism (revkin et al. 2008). counting is effortful, error-prone and slow (trick 1992). it usually takes 250-350 ms per item. in contrast, subitizing is effortless, accurate and rapid. it usually takes 40-100 ms per item. subitizing has been argued to be a preattentive proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 300 https://doi.org/10.3765/elm https://www.elm-conference.net/ mechanism which allows an observer to grasp the cardinality of items without carefully counting them (trick 1992). 2. experiment 1. 2.1. hypotheses. hypothesis 1 was that overspecification of small cardinalities would not differ from overspecification of color. hypothesis 2 was that overspecification of small cardinalities and overspecification of color would be significantly different from overspecification of large cardinalities. 2.2. participants. 90 people took part in the experiment (56 females, age range = 17-32, mean age = 21). 2.3. materials. the experiment had a between-subjects design. in order to verify hypotheses 1 and 2, we created three conditions: color, small cardinality and large cardinality conditions, see figure 1. in all the three conditions, we used 2x2 pictures/cells presented in one slide, each of which contained various geometric objects (squares, rectangles, crosses, circles, triangles, diamonds, stars, and ovals). the geometric objects were identical within each cell but were different among all cells. in all the conditions, one cell was a target and was highlighted, whereas the other three were distractors. a critical cell took different positions through the experiment: it could be any of the 2 x 2 cells. figure 1: examples of critical items used in color, small cardinality and large cardinality conditions (experiment 1) in color condition, two (out of four) cells comprised objects of one color, whereas two other cells comprised objects of another color. there were three colors: red, green and yellow. in small cardinality condition, two (out of four) cells included objects of one small cardinality, while two other cells included objects of another small cardinality. there were three small cardinalities: two, three and four. in large cardinality condition, two (out of four) cells had objects of one large cardinality, whilst two other cells had objects of another large cardinality. there were four large cardinalities: 7 x 8 (56), 8 x 7 (56), 8 x 8 (64). each condition contained 48 critical items, that is, 8 geometric objects x 3 colors x 2 versions in color condition, 8 geometric objects x 3 small cardinalities x 2 versions in small cardinality condition, 8 geometric objects x 3 large cardinalities x 2 versions in large cardinality condition. depending on a condition, we expected that the participants would use the following ways of referring to objects. in color condition, there might be either a singular noun (e.g., “a square”) or a color adjective and a noun (e.g., “a red square”). in small cardinality condition, the options are either a bare plural noun (e.g., “squares”) or a numeral and a plural noun (e.g., “two squares”). in large cardinality condition, there might be either a bare plural noun (e.g., “squares”), a quantifier and a noun (e.g., “many squares”) or a numeral and a noun (e.g., “56 squares”). importantly, a proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 301 https://doi.org/10.3765/elm https://www.elm-conference.net/ singular noun and a bare plural noun are minimal specifications of referring to given highlighted cells. a color adjective and a noun as well as a numeral and a noun are overspecifications. as for combinations of a quantifier and a noun, they do not seem to be minimal specifications. however, they do not seem to be overspecifications either, since adding information of a large amount of some set does not tell anything about a cardinality of such a set, especially if other sets presented in a slide can also be referred to with help of “many” expression. therefore, it seemed reasonable to treat data with “many” (if they would occur) as minimal specifications. filler items were images of human faces, tangrams, and artifacts (crockery, furniture, transport and clothes), cf. figure 2. in parallel to critical items, each filler slide contained four images of two sorts: e.g., two artifacts and two tangrams, two human faces and two artifacts, two tangrams and two human faces. there were 72 filler items that were identical in all the three conditions. figure 2: examples of filler items (experiment 1) the idea behind using human faces as fillers was that their description involves more participants’ concentration because they contain many features important to distinguish one face from others (see koolen et al. 2013): with vs. without beard, with vs. without glasses, hair style, mood of a person, a dress, etc. tangrams seem to be even more difficult than human faces in reference production. on the contrary, artifacts are easier to be referred to, since their identification is effortless. all the fillers were intentionally left uncolored, that is black-and-white. the idea behind that was that participants were not supposed to concentrate on color of hair, dress, etc. 2.4. procedure. there were two versions of each condition with a random order of critical and filler items. importantly, critical and filler items were counterbalanced, so that each pair of critical items was separated with at least one filler. therefore, each condition formed two experimental lists. each list had 120 items (48 critical items + 72 fillers). each list was presented for 30 participants (30 participants for a condition x 3 conditions = 90 participants). the experiment was conducted in the russian language. before the experiment, participants were told that they had to describe a highlighted picture to a person who had the same set of pictures but in a different order and who did not know which cell was highlighted. participants’ task was to describe a highlighted picture to a person so that she understood which cell is referred to. importantly, after pressing the space key or the right arrow on the keyboard, participants moved from one slide to another one. in the instructions, they were asked to make sure that their interlocutor also changed their slide. only when the interlocutor confirmed this, a participant could start describing the picture. this was done intentionally to provide participants with some time to carefully examine a given slide. there were a few practice trials (identical to fillers) before the experiment. because of covid-19, the experiment was conducted online, via zoom. participants gave permission to be audio-recorded. the participants were told that no correct answers are expected. they were instructed not to think too long but not to hurry up. proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 302 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.5. results. 4320 responses (48 critical items x 30 participants x 3 conditions) were received (1440 responses for each condition). however, some of them were excluded due to participants’ metaphorical naming objects, mostly in color condition (e.g., “yellow oval” was described as an antispasmodic pill). out of 1440 responses in color condition, 49 responses (3.4%) were excluded. out of 1440 responses in small cardinality condition, 4 responses (0.28%) were excluded. out of 1440 responses in large cardinality condition, 22 responses (1.5%) were excluded. therefore, 1391 responses for color condition, 1436 responses for small cardinality condition, 1418 responses for large cardinality condition were used. in large cardinality condition, 308 responses (out of 1440, 21%) that specified a large amount of objects without a numeral (e.g., “many ovals”) were treated on a par with bare plurals (e.g., “ovals”). in small cardinality condition, 2 responses (out of 1440, 0.14%) that specified a small amount of objects without a numeral (e.g., “some ovals”) were treated on a part with bare plurals (e.g., “ovals”). the results of experiment 1 are visualized in figure 3. figure 3: distribution of (over)specification in color, small cardinality and large cardinality conditions (experiment 1) a kruskal-wallis rank sum test indicated a main effect of modifier (color vs. small cardinality vs. large cardinality): h (2) = 2642.2, p < 0.0001. multiple pairwise comparisons using wilcoxon rank sum test showed further differences between conditions: color vs. large cardinality (p < .0001), color vs. small cardinality (p < .0001), small cardinality vs. large cardinality (p < .0001). 2.5. discussion. the results of experiment 1 confirmed hypothesis 2 but disconfirmed hypothesis 1. both small cardinalities (2, 3, 4) and color adjectives were overspecified in reference production to a greater extent than large cardinalities. in this respect, small cardinalities and color resemble each other. a plausible reason for this resemblance might be their absoluteness and salience. however, they were overspecified in a different way, with the proportions for small cardinalities higher than the proportions for color adjectives. this is an unexpected finding in the research field of overspecification in reference production that has mostly concentrated on color and tentatively concluded that it is the most overspecified attribute. if small cardinalities are absolute and salient, why are they so? a possible reason is that they undergo the subitizing effect. this is what we tested in experiment 2, where the slides were presented in a flashed mode. additionally, this might be relevant for metaphors that occurred in the participants’ responses. metaphors have been argued to be time-consuming (cf. noveck et al. 0% 20% 40% 60% 80% 100% сolor large cardinality small cardinality no overspecification overspecification proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 303 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2001 among others) in comprehension studies. metaphorical interpretations require longer reaction times than literal interpretations. what about speeded reference production? would they occur if participants had timings to refer to a flashed picture? there is one more question. it has been argued that participants demonstrate consistency in reference strategies (cf. tarenskeen et al. 2015). for example, if they start using a color adjective, they do so almost through the whole experiment. would they do so when they were presented with a flashed picture and when they had timings to refer to it? all these questions are addressed in experiment 2. 3. experiment 2. 3.1. hypotheses. hypothesis 3 was that overspecification of small cardinalities presented in a flashed mode would be similar to overspecification of small cardinalities in experiment 1. hypothesis 4 was that there would be no metaphorical expressions produced while referring to objects. hypothesis 5 was that proportions of (over)specification would be consistent through the whole experiment. 3.2. participants. 31 people took part in the experiment (23 females, age range = 20-30, mean age = 24). 3.3. materials and procedure. both critical and filler items were identical to the items used in the small cardinality condition of experiment 1. however, the procedure was different. the slides for small cardinality condition of experiment 1 were presented in a flashed mode on the screen, cf. figure 4. firstly, participants saw a blank slide with a fixation dot for 500 ms. it was followed by a slide with four cells. each cell contained a unique set of geometric objects. the cardinalities of two cells were identical and the cardinalities of two other cells were also identical. no cell was highlighted. such a slide appeared on the screen for 5 000 ms. during this time interval, participants could carefully examine a given slide. after that, participants were presented with the same slide, however, importantly, one of the cells was highlighted. the presentation of such a slide was for a quite short time, only for 200 ms. the reason for this short time interval was that according to trick (1992) subitizing usually takes 40-100 ms per item, that is on average 70 ms per item. therefore, the interval would be enough to subitize the cardinality in the highlighted cell. a slide presented for 200 ms was followed by a blank slide appeared on the screen for 5 000 ms. during this slide, participants had to describe the highlighted cell of the previous slide. participants were instructed to get ready while being presented with the first slide, to carefully examine four pictures in the second slide, to catch sight of which cell was highlighted in the third slide, and to describe a highlighted cell while being presented with the fourth slide. participants had to describe highlighted cells to a person who had the same set of pictures but in a different order and who did not know which cell was highlighted. that is, instead of a video presentation, a person had a mere presentation (as in experiment 1). there were a few practice trials (identical to fillers) before the experiment. because of covid-19, the experiment was held online, via zoom. participants gave permission to be audio-recorded. the participants were told that there were no correct answers. they were instructed not to think too long but not to hurry up. the video presentation lasted 22 minutes. proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 304 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 4: example of the presentation sequence of a critical item (experiment 2) 3.4. results. 1488 responses (48 critical items x 31 participants) were received. however, 90 responses (6.05%) of them were excluded because of problems similar to those occurred in experiment 1: participants’ self-corrections in counting objects (e.g., “three… four stars”), selfcorrections in specifying cardinalities (e.g., “squares… three squares”), errors in counting objects (e.g., “five ovals” when a cell with four ovals was highlighted). interestingly, there were no metaphorical naming of objects. moreover, there were some technical problems (unstable internet connection) while presenting the materials to the participants via zoom. this fact disallowed to record audio responses to some of the critical items. the retained 1398 responses were used. the results of experiments 1 and 2 are visualized in figure 5. 500 ms 5000 ms 200 ms 5000 ms proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 305 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 5: distribution of (over)specification in color, small cardinality and video small cardinality conditions (experiments 1 and 2) a kruskal-wallis rank sum test showed a main effect of condition (color vs. small cardinality vs. video small cardinality) in experiments 1 and 2: h (2) = 193.17, p < 0.0001. multiple pairwise comparisons using wilcoxon rank sum test revealed further differences between conditions: color vs. video small cardinality (p < .0001), small cardinality vs. video small cardinality (p < .026), color vs. small cardinality (p < .0001). additionally, we tested whether participants were tired during the experiment and whether it affected the results. we calculated proportions of (over)specification in the first part (first 24 critical items) and the second part (next 24 critical items) of the experiment. the results are visualized in figure 6. figure 6: distribution of (over)specification in video small cardinality condition between the first and second part of the experiment (experiment 2) a wilcoxon rank sum test with continuity correction showed a higher proportion of overspecification in the first part of experiment 2: w = 258880, p < 0.0001. 3.5. discussion. due to the significant difference between small cardinality vs. video small cardinality conditions, hypothesis 3 was not confirmed. however, the proportions of overspecification in both cases are visually quite similar and distinct from color condition (cf. figure 5). a plausible reason for this might lie in the mode of presentation. firstly, the experiment con0% 20% 40% 60% 80% 100% color small cardinality video small cardinality no overspecification overspecification 0% 20% 40% 60% 80% 100% first part second part no overspecification overspecification proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 306 https://doi.org/10.3765/elm https://www.elm-conference.net/ tained too many slides (120) and, secondly, the timing of 200 ms for the presentation of a slide with the highlighted was relatively short, even though the previous slide was shown for 5 000 ms. be that as it may, the proportion of overspecification in experiment 2 is still relatively low. this suggests that subitizing plays a role in overspecifying small cardinalities. subitizing makes small cardinalities salient, and, in this respect, they resemble color that has been argued to be also salient (brown-schmidt and konopka 2011, tarenskeen et al. 2015 among others). moreover, in both experiments, numerals were produced in exact meanings ‘exactly n’ (cf. the discussion of which meanings are primary for numerals: at-least meanings ‘at least n and possibly n+1’ vs. exact meanings ‘exactly n’ – in papafragou and musolino 2003, musolino 2004, breheny 2008 as well as more recent studies). a possible explanation for this is again subitizing. it seems natural to assume that subitizing yields the exact meanings of the numerals. this finding has the following consequence related to color. like color, small cardinalities is absolute. hypothesis 4 was confirmed. there was no metaphorical naming of objects in experiment 2. it accords with the idea that metaphors are time-consuming and are not produced in reference production under time pressure. hypothesis 5 was not confirmed. it suggests that consistency is not appropriate under time pressure. 4. general discussion. the two experiments reported in this paper demonstrated that small cardinalities (till 4) are overspecified in reference production because of the two factors: absoluteness and salience. due to these factors, small cardinalities resemble color, which has been argued to be absolute and salient (tarenskeen et al. 2015 among many others). a possible explanation for the salience and absoluteness of small cardinalities is subitizing which has been observed for small cardinalities till 4 (kaufman et al. 1949 among many other studies). subitizing makes small cardinalities salient and forces the corresponding numerals to be used in exact meanings (‘exactly n’) rather than in at-least meanings (‘at least n’). the present study supports the idea that both absoluteness and salience play a crucial role in overspecification of cognitive domains in reference production. additionally, the paper has implications for producing metaphor and inconsistency in reference under time pressure. references akin, omer & william chase. 1978. classification of three-dimensional structures. journal of experimental psychology: human perception and perfomance 4(3). 397–410. arts, anja, maes, alfons, noordman, leo & carel jansen. 2011. overspecification in written instruction. linguistics 49. 555–574. atkinson, janette, campbell, fergus & marcus francis. 1976. the magic number 4 ± 0: a new look at visual numerosity judgements. perception 5. 327–334. belke, eva & antje meyer. 2002. tracking the time course of multidimensional stimulus discrimination: analyses of viewing patterns and processing times during “same”-“different” decisions. european journal of cognitive psychology 14. 237–266. bourdon, benjamin. 1908. sur le temps nécessaire pour nommer les nombres. revue philosophique de la france et de l’etranger 65. 426–431. breheny, richard. 2008. a new look at the semantics and pragmatics of numerically quantified noun phrases. journal of semantics 25(2). 93–139. proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 307 https://doi.org/10.3765/elm https://www.elm-conference.net/ brown-schmidt, sarah & agnieszka konopka. 2011. experimental approaches to referential domains and the on-line processing of referring expressions in unscripted conversations. information 2. 302–326. chi, michelene & david klahr. 1975. span and rate of apprehension in children and adults. journal of experimental child psychology 19. 434–439. gatt, albert, van der sluis, ielka, & emiel krahmer. 2007. evaluating algorithms for the generation of referring expressions using a balanced corpus. proceedings of the 11th european workshop on natural language generation. 49–56. gatt, albert, krahmer, emiel, van gompel, roger & kees van deemter. 2013. production of referring expressions: preference trumps discrimination. proceedings of the 35th annual conference of the cognitive science society. 483–488. kaufman, e., lord, m. & t. reese. 1949. the discrimination of visual number. american journal of psychology 62. 498–525. koolen, ruud, gatt, albert, goudbeek, martijn, & emiel krahmer. 2011. factors causing overspecification in definite descriptions. journal of pragmatics 43. 3231–3250. jensen, e., reese, e. & t. reese. 1950. the subitizing and counting of visually presented fields of dots. journal of psychology 30. 363–392. livingstone, margaret & david hubel. 1988. segregation of form, color, movement, and depth: anatomy, physiology, and perception. science 240. 740–749. mandler, george & billie shebo. 1982. subitizing: an analysis of its component processes. journal of experimental psychology: general 111(1). 1–22. mangold, roland & rupert pobel. 1988. informativeness and instrumentality in referential communication. journal of language and social psychology 7. 181–191. musolino, julien. 2004. the semantics and acquisition of number words: integrating linguistic and developmental perspectives. cognition 93(1). 1–41. noveck, ira, maryse bianco, & alain castry. 2001. the costs and benefits of metaphor. metaphor and symbol 16. 109–121. oyama, tadasu, kikuchi, tadashi & shigeru ichihara. 1981. span of attention, backward masking and reaction time. perception and psychophysics 29(2). 106–112. papafragou, anna, & julien musolino. 2003. scalar implicatures: experiments at the syntaxsemantics interface. cognition 86. 253–282. pechmann, thomas. 1989. incremental speech production and referential overspecification. linguistics 27. 89–110. revkin, susannah, piazza, manuela, izard, véronique, cohen, laurent & stanislas dehaene. 2008. does subitizing reflect numerical estimation? psychological science 19(6). 607–614. rubio-fernández, paula. 2016. how redundant are redundant color adjectives? an efficiencybased analysis of color overspecification. frontiers in psychology 7. 153. sedivy, julie. 2003. pragmatic versus form-based accounts of referential contrast: evidence for effects of informativity expectations. journal of psycholinguistic research 32. 3–23. stevens, stanley. 1938. on the problem scales for the measurement of psychological magnitudes. journal of unified science 9. 94–99. tarenskeen, sammie, broersma, mirjam, & bart geurts. 2015. overspecification of color, pattern, and size: salience, absoluteness, and consistency. frontiers in psychology 6. 1703. taves, ernest. 1941. two mechanisms for the perception of visual numerousness. psychology 37. 1–47. proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 308 https://doi.org/10.3765/elm https://www.elm-conference.net/ thomas, roger, phillips, julia & cheryl young. 1999. comparative cognition: human numerousness judgments. the american journal of psychology 112 (2). 215–233. trick, lana. 1992. a theory of enumeration that grows out of a general theory of vision: subitizing, counting, and finsts. in j. campbell (ed.), the nature and origins of mathematical skills, 257–299. elsevier science publishers b.v. van gompel, roger, van deemter, kees, gatt, albert, snoeren, rick, & emiel krahmer. 2019. conceptualization in reference production: probabilistic modeling and experimental testing. psychological review 126(3). 345–373. westerbeek, hans, koolen, ruud, & alfons maes. 2015. stored object knowledge and the production of referring expressions: the case of color typicality. frontiers in psychology 6. 935. proceedings of elm 1: 298-309, 2021 natalia zevakhina and elena pasalskaya: overspecification of small cardinalities in reference production. 309 https://doi.org/10.3765/elm https://www.elm-conference.net/ (not) acquiring meaning in a second language: are input deficits key? barbara c. malt, xingjian yang, and jessica joseph1* abstract. word meanings are not always parallel across languages, and second language (l2) learners often use words in non-native ways. is the learning problem inherent in maintaining conflicting word-to-meaning mappings within an integrated lexical network, or is it due to insufficient attention to and input for acquiring l2 mappings? to help discriminate between these possibilities, we gave english speakers repeated exposures to 40 brief videos of actions, labeled with five novel words that cross-cut english labeling patterns. half the participants were told only to learn the labels for the actions. the other half were told to figure out their meanings, which might differ from english. the figure out meanings group made test choices faster and were also slightly more likely to produce definitions capturing the intended meanings. however, both groups performed well above chance in generalizing the novel words. high levels of choice performance for both groups point to insufficient input, rather than inherent properties of lexical networks, as the critical limiting factor in more typical l2 learning contexts. speed and definition performance hint at some advantage to explicit attention in sorting out l1-l2 differences. keywords. word learning; word meaning; vocabulary; second language learning; l1 influence on l2 1. introduction. languages are diverse in their lexicons as well as their grammar and sound systems. for instance, english labels a ball on a table and a handle on a door both as on, whereas dutch labels them separately as op vs. aan. on the other hand, spanish labels both along with an apple in a bowl as en (bowerman 1996). languages also vary in how many color terms they use to divide up the spectrum, what distinctions are drawn among drinking vessels, and the meaning of body part terms, among others (for review, see malt & majid 2013). these differences are not easily mastered by second language (l2) learners. they often use words in non-native ways even after many years of l2 immersion (de groot 2012, malt & sloman 2003, pavlenko 2009). the challenge for the l2 learner entails acquiring lexical categories that may only partially overlap with first language (l1) categories, may be subsets or supersets of l1 categories, or may cross-cut them by using entirely different semantic dimensions. in lab studies of artificial category learning, learning one set of categories followed by learning a cross-cutting set produces large costs in speed and accuracy of categorization and perseverative errors (e.g., kruschke 1996). in l2 acquisition in natural (non-laboratory) contexts, cross-cutting also proves difficult. for instance, english speakers carry or hold an object, regardless of manner of support or contact, whereas in korean, chinese, and japanese, the same actions are labeled according to manner of contact or support, regardless of movement (saji, imai, saalbach, zhang, shu & okada 2011). malt (2020) studied native mandarin speakers’ use of english carry and hold and found that even mandarin speakers immersed in an english language environment showed persistent errors in the use of these terms (see also jessen & cadierno 2013, malt, li, pavlenko, zhu & ameel 2015). 1* we thank ashley veliz for assistance with data collection and coding and helen borchart and yichen shen for serving video models. authors: barbara malt, lehigh university (bcm0@lehigh.edu), xingjian yang, lehigh university (xiy221@lehigh.edu), jessica joseph, lehigh university (jlj217@lehigh.edu) proceedings of elm 1: 204-211, 2021 c©2021 barbara c. malt, xingjian yang, and jessica joseph published by the lsa with permission of the author(s) under a cc by license. 204 https://doi.org/10.3765/elm https://www.elm-conference.net/ such results suggest that entrenchment in an initial set of categories interferes with encoding alternative groupings of the same entities (for discussion of entrenchment in language learning, see, e.g., macwhinney 2017, schmid 2017). in lexical network terms, once the network settles into a stable configuration of links between word forms and elements of meaning, remapping may become difficult. furthermore, continued use of the l1 may create a continued pull away from the l2 form-to-meaning mappings (malt et al., 2015). paradoxically, though, within the l1, language users easily master words having complex relations among their meanings. for instance, cup both overlaps with mug and can be used to encompass mugs, and goal-derived and thematic terms cross-cut taxonomic ones (e.g., pet and wildlife crosscut feline and canine; similarly for breakfast foods and dinner foods vs. dairy and grains or things to take on a picnic vs. drinks and utensils (barsalou 1991, ross & murphy 1999). l1 learners acquire and maintain many such cross-cutting terms without difficulty. these observations indicate that at least within a single language, there is not a fundamental difficulty in simultaneously representing different ways of grouping the same entities by name. the contrast between ease of acquiring cross-cutting mappings in l1 and persistent difficulties in l2 raises the question of where the difference lies. is the l2 problem somehow inherent in how the lexical network deals with conflicting l1 and l2 mappings? or could it be due to differences in how l1 and l2 are acquired? young l1 learners often are given many exposures to examples of the same word-referent pairs, and the input often encourages them to focus on extracting meaning (e.g., look, i have a lion. look at the lion. do you know what a lion is?; snow, 1972). both of these input features may facilitate extraction of meaning. l2 learners, especially outside of a classroom, are likely to hear much more diverse input and to need to focus on understanding the larger discourse goals. child l1 learners also have metalinguistic knowledge that meanings are contrastive (clark, 1987; that is, that kitty and lion are different, and that pet and wildlife should differ from feline and canine), whereas l2 learners may simply assume that l2 meanings are parallel to l1 meanings. this is likely true in both formal instruction and in immersion contexts. the current study was designed to discriminate between the possibility that learning difficulty is due to fundamental characteristics of lexical networks and the possibility that it is due to insufficient input and attention to the input. it also investigated whether metalinguistic knowledge that l2 meanings can differ is crucial to successful learning. we used a domain where word meanings cross-cut each other in the l1 and the to-be-learned language. we gave participants intensive l2 word-referent pairing exposure and tested their ability to define the l2 words, label exposed instances, and generalize to new instances. further, half the participants were told simply to learn the l2 labels and half were told that meanings may differ from l1, thereby cueing them to search for new meanings. 2. method. 2.1. participants. seventy-eight native english speakers without substantial knowledge of another language participated. they were recruited from lehigh university’s introductory psychology participant pool. native language and other language experience was assessed with a brief language history questionnaire. data from an additional five participants with exposure to asian languages, arabic, and urdu were not included. 2.2. stimuli. stimuli were designed to have meanings cross-cutting the meanings of english words for them. training stimuli were videos illustrating five novel verbs for actions of standing or walking with an object, modeled after saji et al. (2011). the actions are labeled as carry or hold in english, as established in a pre-test. in many asian languages and in our lab version, these proceedings of elm 1: 204-211, 2021 barbara c. malt, xingjian yang, and jessica joseph: (not) acquiring meaning in a second language: are input deficits key?. 205 https://doi.org/10.3765/elm https://www.elm-conference.net/ actions are labeled according to the manner of the object being in contact with or supported by the person (e.g., cupped in one or both palms vs. held snug against the front or side of the body), regardless of whether the action is stationary or moving (see saji et al. 2011). we created 40 training videos to illustrate 5 novel verbs of interacting with objects. the verbs were based on korean verbs for this domain (as described in saji & imai 2013). we chose the 5 korean verbs instead of the 7 japanese or 13 mandarin verbs described by saji and imai (2013) to limit the memory burden. the real korean meanings varied in their breadth and overlap, but we created meanings of approximately the same scope per verb and without overlap. each verb encoded a manner of contact or support (antta: support on shoulder or back; ida: support on one or two palms; kidda: hug or clutch against front or side; meda: dangle from arm, finger, or hand; teulda: grasp or grip with one or two hands; see figure 1). each of the five novel verbs was shown once as stationary and once with forward movement with each of four objects (yielding 8 video examples per verb and 40 videos total). figure 1: examples of the five verbs to be learned. from left: antta, ida, kidda, meda, teulda test stimuli were new instances of the same five verbs (with four new objects per verb and modest variations in their placement), shown once as stationary and once with movement for each object (yielding 40 test videos). 2.3. procedure. participants were told that they would see brief video clips of a person interacting with an object and that each one would be labeled with a korean verb. the learn labels group was told further only that “you'll try to learn the correct label for each type of action.” the figure out meanings group was told to try to learn the correct label “by figuring out what the verbs mean.” they were also told “note that verbs in korean may have different meanings from the ones you would use in english. keep an open mind about what the meanings of the labels might be.” there were two learning phases and a test phase. in learning phase 1, participants first viewed all the training stimuli with labels once, and then wrote what they thought each label meant. in learning phase 2, they chose a label for each stimulus and received feedback with the correct label. they then wrote again what they thought each label meant. in the test phase, they saw each novel stimulus and chose a label for it from a list of the five. 3. results. we first discuss differences between the groups and then consider implications of the overall performance level of both groups. 3.1. group differences. 3.1.1. latencies. if the learn labels group has more difficulty shifting semantic dimensions and inducing the meanings of the verbs relative to the figure out meaning group, they might take longer to complete the study. they might also take longer to make label choices in the test phase. (there were no choices in training phase 1, and choice times were not recorded in training phase 2). the learn labels group was slightly but not significantly slower than the figure out meanings group to complete the study (t(76) = -1.18, p > .10). the non-significance may be due to the fact that fixed-time portions of the study such as video exposures swamped the impact of performance proceedings of elm 1: 204-211, 2021 barbara c. malt, xingjian yang, and jessica joseph: (not) acquiring meaning in a second language: are input deficits key?. 206 https://doi.org/10.3765/elm https://www.elm-conference.net/ differences in overall time. however, as expected, this group was significantly slower to complete the test trials (t(76) = -1.84, p < .05). (see table 1.) group study completion time test trial choices learn labels 24.84 mins 3.46 sec figure out meanings 21.43 mins 3.10 sec table 1: latencies 3.1.2. definitions. if the learn labels group has more difficulty shifting semantic dimensions and inducing the verb meanings, their definitions should less well capture the intended meaning of the verbs, and especially so in learning phase 1. we scored definitions in two ways: first, based on whether the definitions captured the general idea that meanings were based on the manner of contact or support, and second, whether they captured the specific intended meaning. each definition (for each method) was assigned a 1 if closely capturing the target meaning, 0.5 if partially capturing it, and 0 if not similar to the target. scoring was carried out by the three authors and a research assistant, with blocks of responses arranged in random order so that scorers were blind to condition. each response was scored by all four scorers and its mean score was obtained. in learning phase 1, the learn labels group was slightly but not significantly worse than the figure out meanings group at both producing definitions mentioning manner of contact or support and at capturing the specific intended meaning (t(76) < 1 for both; ps > .10). in learning phase 2, the two groups were virtually identical for both. (see table 2.) group general meaning specific meaning learning phase 1 learn labels .67 .47 figure out meanings .72 .51 learning phase 2 learn labels .79 .58 figure out meanings .77 .57 table 2: mean score, definitions 3.1.3. label choices. if the learn labels group has more difficulty shifting semantic dimensions and inducing the meanings of the verbs relative to the figure out meaning group, their choices of labels for the actions should be less accurate in both learning phase 2 and in the test phase. (no choices were made in learning phase 1). the groups did not differ in learning phase 2, and, contrary to expectation, the learn labels group did somewhat better in the test phase (t(76) = 1.6, p = .06; see table 3). 3.2. overall performance. overall, both groups performed at high and similar levels and showed improvement over the course of exposure. 3.2.1. definitions. as seen in table 2, each group increased the extent to which their definitions captured the general idea of body contact or support from the 1st to the second round of definitions, and each group also increased the extent to which their definitions captured the specific intended meaning from the first to the second round of definitions. the changes were significant for the learn labels group (t(70) = 1.84 for general idea and t(70) = 1.73 for specific intended meaning, p < .05 for both) but not for the figure out meanings group, whose performance was higher to begin with (for both, t(82) < 1.0, ps > .10). proceedings of elm 1: 204-211, 2021 barbara c. malt, xingjian yang, and jessica joseph: (not) acquiring meaning in a second language: are input deficits key?. 207 https://doi.org/10.3765/elm https://www.elm-conference.net/ group % correct learning phase 2 learn labels .74 figure out meanings .73 test trials learn labels .83 figure out meanings .75 table 3: label choices 3.2.2. label choices. because there were five labels to choose from for each stimulus, chance performance would be 20% correct. as seen in table 3, performances of both groups at learning phase 2 and at the test phase were well above this level (tested against .20, all ts > 12, all ps < .0001), indicating that substantial learning took place. 4. discussion. l2 learners often show difficulty reshaping l1 word meanings to l2 meanings. is this difficulty inherent in the nature of lexical networks (due to l1 entrenchment and competition between mappings), or is it due to the nature of typical l2 learning conditions? our study was designed to help discriminate between these possibilities. it provided intensive exposure to novel word-referent pairings, fostering abstraction of commonalities and contrasts among the words and facilitating a focus on word meaning extraction rather than on understanding a larger discourse. half the participants were also cued to expect that meanings might differ from l1. 4.1. group comparison. half of our participants were asked simply to learn the novel labels for the actions they viewed, and half were cued to look for meanings different from those of their l1. the figure out meanings group was significantly faster to make choices in the final test. they also had a non-significant edge in producing correct definitions of the novel words after the first round of exposure. however, they did not have more correct choices of labels in the second learning phase or the test phase. because there was no hint of advantage for the figure out meanings group and the sample sizes were not small, it is unlikely that stronger evidence of choice differences would emerge with additional participants. the similar choice performance between groups indicates that learners can do well even without prior knowledge that they might need to attend to different dimensions. the fact that their native language provides only two labels for the actions (carry and hold) whereas they were asked to learn to discriminate among use of five different korean labels may have been sufficient to alert both groups that they needed to approach the task looking for meanings different from carry and hold. in real world l2 exposure, learners may not encounter the complete vocabulary for a domain within a single context, making it much less obvious that meaning differences must exist. 4.2. overall performance. both groups produced levels of choice performance well above choice in the test phase, despite the short study duration. this outcome indicates that it was not difficult to transition from their l1 system to a different system of labeling, again favoring the possibility that it is typical conditions of learning rather than inherent properties of lexical networks that create difficulty in real world l2 acquisition. however, it should be noted that performance on the final test did not exceed 83% correct, and it will be important to examine the nature of the errors in future analyses, in conjunction with examining errors in definitions produced at the second learning phase. it was apparent in scoring the definitions that learners sometimes simply confused which label went with which meaning proceedings of elm 1: 204-211, 2021 barbara c. malt, xingjian yang, and jessica joseph: (not) acquiring meaning in a second language: are input deficits key?. 208 https://doi.org/10.3765/elm https://www.elm-conference.net/ (although they had induced relevant meanings). there was also some evidence in the definitions that a minority of participants did have difficulty making the meaning transition, persisting in producing definitions that referred to moving or standing still with the objects. it remains to be determined what the bulk of label choice errors in the test phase consist of, and whether those errors point to a particular source of difficulty in transitioning from english meanings to the novel verb meanings. 4.3. additional considerations. aside from input conditions, the task given to our participants was easier than the task that would be facing english learners of real korean (or mandarin or japanese) in several respects. we designed our novel verb meanings to have approximately equal scope and no overlap among the meanings. in reality, these languages have terms with fuzzy boundaries and some that are broad in scope and others that are more specific. we also limited the number of verbs to five, whereas some languages have up to about 13 terms regularly used within the domain (saji & imai 2013). such variations are likely to provide extra challenges to the learner in terms of figuring out the meanings associated with each term. however, although inducing the correct meanings may be more difficult under these circumstances, understanding that the nature of the meanings must fundamentally differ from that of the english system should, if anything, be even more clear. another consideration is that this work was prompted by observations of difficulty by l1 mandarin speakers in grasping english carry and hold (malt, 2020), but (to obtain larger sample sizes) the current study taught l1 english speakers meanings modeled on the asian languages. thus, the participants were moving from a simpler, two-word system to a more complex five-word system. would learning be different if we had asked speakers of an asian language to learn english meanings in the lab task, where they would need to collapse multiple meanings to fewer instead of vice versa, as well as make a dimensional shift? previous work on learning in natural contexts suggests that going from more to fewer meanings is not necessarily easier than going from few to many (e.g., jessen & cadierno 2013, gathercole & moawad 2010). ultimately, however, this is an empirical question, and we are currently preparing to run l1 mandarin speakers in china in such a task. another difference between moving from l1 english to l2 korean-like meanings compared to vice versa is that the korean-like meanings were based on different manners of interacting with objects, whereas the english carry-hold distinction is based on presence or absence of movement. even infants are sensitive to both the manner and path of events (pulverman, song, hirsh‐pasek, pruden & golinkoff 2013). however, some evidence suggests that they may be more sensitive to path (konishi, pruden, golinkoff & hirsh-pasek 2016). to the extent that the english carry-hold distinction is similar to a path distinction, its salience could provide a boost in identifying the relevant underlying dimension for those starting with korean-like (or real korean or japanese or mandarin) meanings. however, the current study also makes clear that noticing the different manners was not difficult under favorable conditions. again, an empirical test is needed to establish if there is a difference in ease of moving from one system to the other. finally, although the current results suggest that learners can shift to new meanings with suitable input, an open question is whether such shifts will be accompanied by alterations to the l1 word representations. some prior research points to the possibility that l2 attainment does exert an influence on l1 form-to-meaning mappings (e.g., malt et al. 2015, pavlenko & jarvis 2002), perhaps by influencing connection strengths between forms and meanings (malt et al. 2015). if so, it will be of interest to pinpoint how properties of lexical networks interact with language input in maintaining or altering representations in the two-language case. proceedings of elm 1: 204-211, 2021 barbara c. malt, xingjian yang, and jessica joseph: (not) acquiring meaning in a second language: are input deficits key?. 209 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4.4. conclusions. high levels of choice performance in our task indicate that most learners can readily pick up on new dimensions of l2 word meaning that cross-cut their familiar l1 meanings. success in the current task context is likely due to viewing many instances of word-referent pairings in succession (fostering abstraction of commonalities per word and identification of contrasts among the words) and attention dedicated to the learning task rather than to other tasks that may be more important in real-world contexts. as such, the results point to input conditions rather than inherent conditions of lexical networks as a critical limiting factor in l2 word learning. references barsalou, lawrence. 1991. deriving categories to achieve goals. psychology of learning & motivation 27. 1–64. bowerman, melissa. 1996. learning how to structure space for language: a cross-linguistic perspective. in paul bloom, mary a. peterson, lynn nadel, and merrill f. garrett (eds.), language and space, 385-436. cambridge, ma: mit press. clark, eve. v. 1987. the principle of contrast: a constraint on language acquisition. in brian macwhinney (ed.), mechanisms of language acquisition, 1–33. hillsdale, nj: lawrence erlbaum associates, inc. groot, annette. m. b. de. 2012. vocabulary learning in bilingual first-language acquisition and late second-language learning. in miriam faust (ed.), the handbook of the neuropsychology of language, vol 1, 472-493. wiley-blackwell, oxford: uk. doi: 10.1002/9781118432501.ch23 gathercole, virginia c. m. & ruba a. moawad. 2010. semantic interaction in early and late bilinguals: all words are not created equally. bilingualism: language and cognition 13. 385408. jessen, moiken, & teresa cadierno. 2013. variation in the categorization of motion in l2 danish by german and turkish native speakers. in juliana goschler & anatol stefanowitsch (eds.), variation and change in the encoding of motion events, 133–162. amsterdam: john benjamins. konishi, haruka, shannon m. pruden, roberta m. golinkoff & kathy hirsh-pasek. 2016. categorization of dynamic realistic motion events: infants form categories of path before manner. journal of experimental child psychology 152. 54-70. kruschke, john k. 1996. dimensional relevance shifts in category learning. connection science 8. 225-247. malt, barbara c. 2020. understanding l2 word learning outcomes: the roles of semantic relations, input, and language dissimilarity. international journal of bilingualism 24. 478–491. malt, barbara c., ping li, aneta pavlenko, huichun zhu & eef ameel. 2015. bidirectional lexical interaction in late immersed mandarin-english bilinguals. journal of memory and language 82. 86-104. malt, barbara c. and asifa majid. 2013. how thought is mapped onto words. wiley interdisciplinary reviews in cognitive science 4. 583-597. doi: 10.1002/wcs.1251. malt, barbara. c. and steven a. sloman. 2003. linguistic diversity and object naming by nonnative speakers of english. bilingualism: language and cognition 6. 47-67. pavlenko, aneta. 2009. conceptual representation in the bilingual lexicon and second language vocabulary learning. in aneta pavlenko (ed.), the bilingual mental lexicon: interdisciplinary approaches, 125 – 160. bristol, u.k.: multilingual matters. proceedings of elm 1: 204-211, 2021 barbara c. malt, xingjian yang, and jessica joseph: (not) acquiring meaning in a second language: are input deficits key?. 210 https://doi.org/10.3765/elm https://www.elm-conference.net/ pavlenko, aneta & scott jarvis. 2002. bidirectional transfer. applied linguistics 23. 190-214. pulverman, rachel, lulu song, kathy hirsh‐pasek, shannon m. pruden & roberta m. golinkoff. 2013. child development 84. 241-252. ross, brian h. and gregory l. murphy. 1999. food for thought: cross-classification and category organization in a complex real-world domain. cognitive psychology 38. 495–553. saji, noburo, mutsumi imai, henrik saalbach, yuping zhang, hua shu & hiroyuki okada. 2011. word learning does not end at fast-mapping: evolution of verb meanings through reorganization of an entire semantic domain. cognition 118. 45–61. saji, noburo & mutsumi imai. 2013. evolution of verb meanings in children and l2 adult learners through reorganization of an entire semantic domain: the case of chinese carry/hold verbs. scientific studies of reading 17. 71-88. snow, catherine e. 1972. mothers’ speech to children learning language. child development 43. 549–565. proceedings of elm 1: 204-211, 2021 barbara c. malt, xingjian yang, and jessica joseph: (not) acquiring meaning in a second language: are input deficits key?. 211 https://doi.org/10.3765/elm https://www.elm-conference.net/ five degrees of (non)sense: investigating the connection between bullshit receptivity and susceptibility to semantic illusions dario paape* abstract. individual differences in people’s tendency to see bullshit statements such as perceptual reality transcends subtle truth as meaningful and possibly profound have become an active topic of research in judgment and decision making in recent years. however, (psycho)linguistics has so far paid little attention to the topic, despite its obvious appeal for language processing research. i present an experiment that investigated possible shared traits contributing to individual bullshit receptivity and susceptibility to semantic illusions, which occur when compositionally incongruous sentences receive plausible but unlicensed interpretations (e.g., more people have been to russia than i have). the results show relatively little indication of an individual-level tendency to both fall for bullshit and for linguistic illusions. implications for future psycholinguistic research into bullshit processing are discussed. keywords. bullshit, semantic illusion, experimental semantics, individual differences 1. introduction. consider the three sentences in (1), which exemplify different types of bullshit or obscurantism (buekens & boudry 2015).1 (1) a. hidden meaning transforms unparalleled abstract beauty. b. a cyclic ionization whose only final result is to induce plasmatic solubility is impossible. c. deliberate ambivalence is inherent to the approach, yielding qualities where things convulse and stutter in emerging vitality. (1-a) is an example “pseudo-profound” or “pseudo-transcendental” bullshit (pennycook et al. 2015, čavojová et al. 2020), (1-b) is an example of “scientific” bullshit (evans et al. 2020), and (1-c) is an example of “international art english” (turpin et al. 2019, rule & levine 2012). the common feature of these three utterances is that they are unclarifiably unclear (cohen 2002): it’s impossible to restate in plain language what the sentences mean, if they mean anything at all. they contain a multitude of abstract terms, jargon, and buzzwords, which are “put together randomly in a sentence that retains syntactic structure” (pennycook et al. 2015; p. 548). bullshit is often produced when speakers talk about things they have no detailed knowledge of, or about things that lack an established “ground truth” (such as art or philosophy; frankfurt 1986). yet, despite its apparent vacuousness, bullshit is sometimes perceived as being meaningful or even profound.2 readers vary in their bullshit receptivity, that is, in their tendency to imbue bullshit with meaning. iacobucci & cicco (in press) present a review of 40 studies published between 2015 and 2021 that are concerned with individual differences in bullshit receptivity. among the *author: dario paape, university of potsdam (paape@uni-potsdam.de). 1buekens & boudry (2015) draw a contrast between bullshit and obscurantism that is based on the utterer’s intention, but under the content-based characterization of bullshit that i will adopt, the contrast becomes irrelevant. 2meaningfulness is, of course, subjective to some extent, and apparent nonsense may be argued to sometimes contain deep wisdom (dalton 2016). however, this is especially unlikely for sentences that are intentionally generated as bullshit (pennycook et al. 2016). proceedings of elm 2: 189-201, 2023 c©2023 dario paape published by the lsa with permission of the author(s) under a cc by license. 189 https://doi.org/10.3765/elm https://www.elm-conference.net/ factors investigated are intelligence (bainbridge et al. 2019), individual proclivity towards analytic versus intuitive thinking (e.g. pennycook et al. 2016), and self-regulation ability (petrocelli et al. 2020). for the present investigation, i will focus on two additional traits: interpretive charity, that is, a person’s tendency to assume that statements are meaningful and true by default (e.g., sperber 2010), and illusory pattern perception or apophenia, that is, a person’s tendency to see patterns where none exist (deyoung et al. 2012). an indiscriminate bias to judge sentences as meaningful and profound has been argued to contribute to pseudo-profound bullshit receptivity in particular (pennycook et al. 2015), though the ability to discriminate between actual profundity and bullshit appears to be a stronger driver of individual differences (bainbridge et al. 2019). apophenia or illusory pattern perception as a driver of bullshit receptivity has been investigated by walker et al. (2019) and by bainbridge et al. (2019). walker et al. (2019) presented participants with random sequences of 10 coin tosses (e.g., hthhtttthh) and asked them to indicate on a 1–7 scale how strongly they felt that the results were random or pre-determined. in another experiment, participants were shown images of objects embedded in visual noise, as well as images containing only visual noise, and were asked whether the image contained an object. profundity ratings for pseudo-profound bullshit were positively correlated both with the coin-toss determinism measure and with the tendency to perceive objects in pure noise. bainbridge et al. (2019) used a more indirect measure of apophenia, by asking participants about various paranormal and otherwise unusual beliefs (“it is possible to move material objects with only one’s thoughts”). paranormal beliefs correlate with illusory pattern perception (van prooijen et al. 2018). bainbridge et al. (2019) found a positive correlation between apophenia and pseudo-profound bullshit receptivity, in line with walker et al.’s findings. to my knowledge, despite interesting implications for semantic and pragmatic processing, bullshit has not received much attention in psycholinguistics.3 the fact that some people find the sentences in (1) meaningful suggests a role for processes that create meaning independently of the compositional makeup of the utterance and the lexical meanings of the words, to the extent that the average reader even has access to the lexical meaning of a word like ionization. however, there is a related class of phenomena that has been studied in psycholinguistics: semantic illusions. a semantic illusion occurs when a semantically incongruous and arguably meaningless sentence appears well-formed and sensible. the sentences in (2) are examples of such illusions. (2) a. more people have been to russia than i have. b. no head injury is too trivial to be ignored. sentence (2-a) is an instance of the so-called comparative illusion (ci). the sentence is compositionally incongruous because an amount of people (more people than x) is compared with an event (i have [been to russia]), which should not yield a sensible meaning. yet, the anomaly of ci sentences is often not consciously detected, and the sentences are rated as relatively acceptable (e.g., o’connor 2015, wellwood et al. 2018). the illusion is not random, however: ci sentences are rated more highly when they contain repeatable rather than unrepeatable events (?more girls 3the only published source i could find in which a link is made is de almeida (2018). proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 190 https://doi.org/10.3765/elm https://www.elm-conference.net/ graduated high school than the boy did), suggesting a preference for an illicit event comparison reading (wellwood et al. 2018, de dios-flores 2016), but also when they contain semantically plural object nps (more cats have mouse toys than the dog does; o’connor 2015), suggesting the availability of an equally illicit cardinality comparison (more mouse toys). in both cases, the ungrammatical sentence is apparently coerced into an interpretable form. sentence (2-b) is a so-called depth charge (dc) sentence (e.g., sanford & sturt 2002). like ci sentences, dc sentences are compositionally incongruous: the degree phrase too trivial to be ignored presupposes that more trivial things are less likely to be ignored, which runs contrary to world knowledge (compare too trivial to be treated). furthermore, (2-b) is usually interpreted to mean don’t ignore head injuries, but the compositional meaning is ignore all head injuries (compare no landmine is too small to be banned; wason & reich 1979). different theories have attributed the dc illusion to superficial processing (wason & reich 1979, paape et al. 2020), to unconscious repair of an assumed speech error (zhang et al. 2022), or to the availability of a stored grammatical construction with an idiosyncratic meaning (fortuin 2014, cook & stevenson 2010). as for ci sentences, mechanisms beyond standard compositional semantics must be recruited in order to make sense of the construction. both ci sentences and dc sentences show considerable variability in judgments between speakers (wellwood et al. 2018, leivada 2020, paape et al. 2020). some speakers tend to almost always find illusion sentences acceptable while others categorically reject them. however, to my knowledge, a possible shared susceptibility to different semantic illusions within the same speaker has never been investigated. nevertheless, it is plausible that such general differences in susceptibility exist. both for ci and for dc sentences, it has been suggested that there is a threshold of processing complexity that differs between speakers, and that processing is aborted once this threshold is reached and the meaning representation is deemed “good enough” (paape et al. 2020, paape 2021, leivada 2020). low threshold settings can result in failures to notice linguistic anomalies and cause acceptability illusions for malformed sentences (e.g., christianson 2016). this perspective is in line with the proposal that readers have individual standards of coherence that affect their depth of processing during reading (van den broek et al. 2011, 2001), with lower standards of coherence meaning less processing. which brings us back to bullshit: it would appear that in order to accept a bullshit statement such as hidden meaning transforms unparalleled abstract beauty as meaningful, one would need to set a relatively low standard of coherence, because the more one thinks about the semantic content of the sentence, the less meaningful it becomes. anecdotally, the same is often true for ci and dc illusion sentences. at the same time, however, both bullshit sentences and illusion sentences require enrichment of the linguistic input to be perceived as meaningful:4 readers presumably want the sentences to make sense, and thus project plausible meanings into them. superficial processing on the one hand does not necessarily contradict semantic enrichment on the other: readers may simply shift their focus away from the actual structure of the input and onto other, stimulus-independent sources of meaning, such as their pre-existing knowledge or opinions (paape 2021). highly apophenic individuals with a “greater tendency to go beyond the available data” (walker et al. 2019; p. 111) and to “creat[e] meaning where no meaning exists” (ibid., p. 117) 4see de almeida (2018) for a similar point. proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 191 https://doi.org/10.3765/elm https://www.elm-conference.net/ may be especially prone to this form of enrichment. the prediction is thus that the perceived meaningfulness of bullshit sentences, ci sentences and dc sentences should covary jointly within speakers, and that highly apophenic readers should be especially likely to see meaning in all three sentence types. but how can individual differences in apophenia be distinguished from differences in interpretive charity – or are they ultimately the same thing? there are reasons to assume that they are not. illusory pattern perception, which is at the core of aponenia, is triggered by cues that at least resemble a pattern (walker et al. 2019). thus, the more a given sentence resembles a wellformed utterance, the more apophenic interpretation should occur. crucially, bullshit sentences and illusion sentences are not nonsense in the sense that they are random jumbles of words.5 on some level, these sentences look and feel “good enough” to pass inspection. interpretive charity, on the other hand, could plausibly make anything appear meaningful: a highly charitable reader may force even “word salad” to have meaning (fowler 1969), just like any artwork can be imbued with meaning by a charitable viewer. in what follows, i present a web-based experiment that investigated possible correlations between bullshit receptivity, susceptibility to semantic illusions, and apophenia. interpretive charity was controlled for by also including sensible sentences and nonsense sentences in the experiment. the experiment used german sentences, and was run with a sample of german native speakers. 2. experimental study. the experiment deviated from previous studies on bullshit by having participants judge the meaningfulness of the sentences instead of their profundity (e.g., pennycook et al. 2015) or truthfulness (e.g., evans et al. 2020). this more basic level of judgment was chosen to allow comparison with the illusion sentences, which are not intended to be profound or necessarily true. three types of bullshit were included: pseudo-profound bullshit, scientific bullshit, and international art english. pseudo-profound bullshit receptivity has been found to correlate both with scientific bullshit receptivity (evans et al. 2020) and with receptivity towards international art english (turpin et al. 2019), suggesting shared underlying traits. in order to somewhat offset participants’ interpretive charity and to naturalize the judgment of unusual sentences, participants were told to imagine that the sentences had been generated by an ai system, and that they were helping to improve the system by distinguishing between “good” and “bad” sentences. the pattern recognition task intended to measure apophenia was also embedded in the ai scenario (see below). 2.1. participants. 100 self-identified german native speakers were recruited on prolific (https: //www.prolific.co; palan & schitter 2018). they were paid £3,50 for their participation.6 5for an illuminating discussion of different kinds of nonsense, see diamond (1981). 6because prolific is based in the uk, renumeration is calculated in british pounds. proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 192 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.2. materials. example items for each experimental condition are shown in (3).7 (3) a. bullshit the invisible is beyond new timelessness. (pseudo-profound) energy can deteriorate based on closed-circuit alliterations of an afocal system. (scientific) the banality of some gestures is disconcerting, and in their strangeness, they convey a future created in the past. (international art english) b. comparative illusion (ci) more spectators in the theater are proud americans than the actor is. c. depth charge illusion (dc) in the end, you realize that no plan is too unrealistic to be scrapped. d. sensible your teacher can open the door, but you must enter by yourself. (profound) newborns require constant attention. (mundane) e. nonsense one can say that flowers with a lot of old nettles do not limp without great experience value. each participant rated 24 bullshit sentences in total (8 of each type), as well as 16 ci sentences and 12 dc sentences.8 these were randomly intermixed with 24 sensible sentences (12 profound, 12 mundane) and 20 nonsense sentences. the nonsense sentences were generated in a stream-ofconsciousness manner by the author and validated as being nonsense by three german speakers. the 20 stimuli for the pattern recognition task were scatterplots created in r (r core team 2022). for each plot, 20 random floating-point numbers between 0 and 30 were generated for both the x and the y coordinate using r’s runif() function. examples are shown in figure 1. participants with high apophenia were expected to “detect” more patterns in the scatterplots, and possibly show longer response times, because they may spend more time trying to find patterns. 2.3. procedure. participants provided informed consent prior to experimentation. the ai scenario was first introduced. participants were told that they should rate the meaningfulness of each sentence in comparison to a completely meaningless baseline sentence (bats don’t go defiantly into the computer next to love without pumping). this baseline sentence was presented in each trial. the rating scale for the target sentences ranged from 1 (“equally bad”) to 7 (“much better”). participants could freely reread the sentences as many times as they wished, but were instructed not to “overanalyze” the sentences and to rely more on their linguistic intuition. reading times in 7pseudo-profound bullshit sentences, scientific bullshit sentences and international art english sentences were translated and adapted from pennycook et al. (2015), evans et al. (2020), and turpin et al. (2019). comparative illusion sentences were translated and adapted from o’connor (2015). depth charge sentences were adapted from paape et al. (2020). sensible sentences were mostly translated and adapted from pennycook et al. (2015), but also featured some new additions. 8a total of 24 dc sentences were created, but only 12 were presented to each participant to limit the duration of the experimental session and to prevent carryover effects. proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 193 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: stimuli used for the pattern recognition task. milliseconds were recorded in addition to the ratings. after all sentences had been rated, the pattern recognition task was administered. participants were presented with the scatterplots in random order and were told that these represented the “neuronal activations” of the ai model, which are sometimes random and sometimes structured. they were instructed to give a binary judgment of whether they intuitively felt that each pattern was structured as opposed to random. participants were also told that all patterns may be random, or all patterns may be structured. 2.4. data analysis. the statistical analysis was carried out using the stan language for bayesian inference (stan development team 2022). because the rating data are ordinal, they were analyzed with a cumulative logit model. this kind of model assumes a continuous latent variable underlying the likert scale ratings, together with a set of thresholds or “cutpoints” that group the latent values into the rating “bins” enforced by the discrete scale (liddell & kruschke 2018). because the individual differences on the latent scale are of interest, and because participants may differ in their use of the discrete scale, recovering the underlying continuous values by modeling participants’ individual cutpoints is of major importance (see below for the implementation). reading times were analyzed assuming a lognormal likelihood. for both dependent variables, a hierarchical model with correlated “random” (or varying) intercepts and slopes by participants and by items was fitted (e.g., pinheiro & bates 2000, gelman & hill 2007). this means that an individual adjustment to the population-level estimate was estimated for each subject and for each item in order to account for the non-independence of measurements and to capture inter-individual differences (barr et al. 2013). in the present study, the correlations between these adjustments, which are also estimated from the data, are of major interest: the question is whether someone who gives higher-than-average ratings to bullshit sentences will, for instance, also give higherthan-average ratings to comparative illusion sentences. in stan, it is possible to set up a model that estimates a separate average (intercept) value for each sentence type, along with participantand item-specific adjustments, and, crucially, correlations between the adjustments across sentence types. a shortened and simplified notation of the model is shown in equation 1. the model code and data are available at https://osf.io/ 54wh7. proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 194 https://doi.org/10.3765/elm https://www.elm-conference.net/ for each trial n, the rating comes from an ordered logistic distribution with a vector c of six cutpoints. the values γ1...6 demarcate the population-level boundaries between rating “bins” in logit space, which are adjusted for each participant in order to account for differences in the use of the likert scale. a population-level intercept α on the latent scale is estimated for each sentence type and adjusted by participant (ui) and by item (wj). the adjustments follow a normal distribution with a to-be-estimated standard deviation σu, σw. the correlations between the adjustments u1...m and w can be recovered from the estimates of the variance-covariance matrix of the random effects. regularizing priors were used for all parameters, including lkj priors (lewandowski et al. 2009) with η set to 2 for the correlations. µ1,n = α1 + u1,i + wj (bullshit) µ2,n = α2 + u2,i + wj (comparative) µ3,n = α3 + u3,i + wj (depth charge) c1,n = γ1 + u4,i c2,n = γ2 + u5,i . . . u1,i ∼ normal(0, σ u1) . . . wj ∼ normal(0, σ w) ratingn ∼ ordered logistic(µx,n, cn) (1) participants’ reading times as well as their reaction times and judgments for the scatterplots were also analyzed within the same model.9 this approach allows for the direct estimation of the critical random-effects correlations across measures, without any pre-aggregration of the data, and thus without any loss of information. furthermore, the addition of item-specific adjustments allows for better estimation of the subject-level individual differences: if a given participant sees a pattern in a scatterplot that has an above-average “pattern-likeness” across all participants, this is much less informative than if the participant sees a pattern in a scatterplot with a below-average “pattern-likeness”. for the reading times, a slope for the number of characters in the sentence was added to the model, along with adjustments by subject. the slope adjustment for each participant was intended as a measure of their tendency towards superficial or “good enough” processing. it has been found that word length effects are reduced during “mindless” reading (schad et al. 2012, reichle et al. 2010, franklin et al. 2011), so it is plausible that the less attention a participant is paying to the sentences, the less the overall sentence length should affect their reading times. 2.5. results. table 1 shows mean ratings and standard deviations of the by-subject means by sentence type. 9it is advantageous for this approach to have the dependent variables on similar scales. here, the binary judgments and the sentence ratings are analyzed on the log-odds (logit) scale while reading times and reaction times are analyzed on the log scale. proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 195 https://doi.org/10.3765/elm https://www.elm-conference.net/ rating sd bullshit 4.37 0.85 comparative 4.42 1.11 depth charge 5.08 0.92 sensible 6.58 0.52 nonsense 2.05 0.77 table 1: mean ratings and standard deviations of by-subject means by sentence type. unsurprisingly, sensible sentences were rated the most meaningful in comparison to the baseline sentence while nonsense sentences were rated the least meaningful. the three other sentence types fall in between, with somewhat higher ratings for dc sentences than for bullshit and ci sentences. ci sentences showed the greatest rating variability between subjects, followed by depth charge sentences. sensible sentences showed the lowest variability. figure 2 shows the correlation estimates for the participant-level random effects extracted from the stan model, along with the associated 95% highest density intervals (hdis) of the posterior distributions (e.g., kruschke 2014) computed using the bayestestr package (makowski et al. 2019). nonsense sensible bullshit depth charge comparative pattern pattern rt length rt -0.5 0.0 0.5 -0.5 0.0 0.5 -0.5 0.0 0.5 -0.5 0.0 0.5 -0.5 0.0 0.5 -0.5 0.0 0.5 -0.5 0.0 0.5 -0.5 0.0 0.5 nonsense sensible bullshit depth charge comparative pattern pattern rt length rt figure 2: 95% hdis of the correlation estimates for the participant-level random effects. the diagonal is not plotted because the correlation of each adjustment with itself is trivially 1. colors show the sign and magnitude of the correlation estimate (red = negative, green = positive). focusing first on the main prediction, namely a positive correlation between apophenia, as measured by increased pattern spotting and longer reaction times for the scatterplots, and the perceived meaningfulness of bullshit sentences and illusion sentences, the data do not show any strong indication of the predicted relationship. only the perceived meaningfulness of comparative illusion sentences shows some indication of a positive correlation with pattern spotting at the participant level (95% hdi: [−0.03, 0.43]). furthermore, the hdis of the pairwise correlations between byparticipant adjustments for bullshit meaningfulness, ci meaningfulness and dc meaningfulness are centered around values close to zero, indicating no evidence for the predicted positive correlation due to individual differences in interpretive charity. proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 196 https://doi.org/10.3765/elm https://www.elm-conference.net/ there is a negative correlation at the participant level between the perceived meaningfulness of sensible sentences and that of nonsense sentences (hdi: [−0.67, −0.19]). perceived meaningfulness of nonsense is negatively correlated with the effect of sentence length on reading times (hdi: [−0.76, −0.36]), while perceived meaningfulness of sensible sentences is positively correlated with the sentence length effect (hdi: [0.27, 0.7]). this suggests that more superficial readers, who were less affected by sentence length, perceived nonsense as more meaningful and sensible sentences as less meaningful than average. superficial or inattentive readers thus seem less able or willing to distinguish between the sentence types, compared to more attentive readers. finally, participants who spent more time looking at the scatterplots showed smaller effects of sentence length on reading time (hdi: [−0.52, −0.08]), as well as higher perceived meaningfulnesss of nonsense sentences (hdi: [−0.03, 0.41]). these correlations can be seen as tentative evidence for a connection between apophenia and superficial language processing. however, there is no indication in the data that individuals who spent more time looking for patterns in the scatterplots ended up “finding” more of them. 3. discussion. the goal of this study was to investigate individual differences in the processing of bullshit statements (hidden meaning transforms unparalleled abstract beauty) and two types of semantic illusion, namely the comparative illusion (ci; more people have been to russia than i have) and the depth charge illusion (dc; no head injury is too trivial to be ignored). the underlying intuition was that readers who find meaning in bullshit statements may also find meaning in semantic illusion sentences. based on the existing bullshit literature, two traits were hypothesized to contribute to both bullshit receptivity and illusion receptivity: the first is a person’s general bias towards finding statements meaningful or even profound, irrespective of content, that is, their tendency towards interpretive charity (e.g. pennycook et al. 2015). the second is a person’s tendency to actively create meaning by seeking patterns for which there is no objective evidence in the data, that is, their tendency towards apophenia (walker et al. 2019). statistically, shared individual differences in interpretive charity, apophenia, and the perceived meaningfulness of bullshit and semantic illusions were explored by looking at the correlations of subject-level random effects in a hierarchical model. however, there was little evidence in the data that would support the hypothesized connections. there was no indication that people who found bullshit sentences more meaningful also found illusion sentences more meaningful, or that finding ci sentences meaningful correlates with finding dc sentences meaningful. regarding effects of apophenia, only the perceived meaningfulness of ci sentences but not that of bullshit sentences or dc sentences showed a positive correlation with finding patterns in random scatterplots, thus casting doubt on a shared underlying mechanism of “meaning creation”. the absence of evidence for a connection between apophenia and bullshit endorsement is unexpected given earlier findings by walker et al. (2019). however, the mismatch may be due to a variety of factors: the smaller participant sample, the use of a mixture of bullshit “genres” in the current study, and/or differences between the patttern-spotting tasks. for instance, unlike the one used by walker et al. (2019), the pattern spotting task used in the current study did not have a baseline in which real patterns needed to be identified. furthermore, the current experiment used judgments of meaningfulness whereas that of walker et al. used judgments of profundity. finally, participants in the current study were explicitly told to rely on their intuition, and some participants proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 197 https://doi.org/10.3765/elm https://www.elm-conference.net/ may have followed this instruction more strictly than others. despite the differences between studies, there was an isolated positive correlation of apophenia with the perceived meaningfulness of ci sentences in the current study, which should be further investigated in future work. interestingly, participants’ reactions to the two additional sentence types tested, namely sensible sentences (your teacher can open the door, but you must enter by yourself ) and nonsense sentences (instead of big fish, you can also rarely flip seven potatoes) did show some clear correlations: people who tended to see more meaning in nonsense tended to see less meaning in sensible sentences, and vice versa. furthermore, this tendency was connected to participants’ depth of processing: the reading times of participants who distinguished less between nonsense and sensible sentences were less affected by sentence length, suggesting that they were paying less attention. in light of this findings, it is not clear why bullshit sentences and illusion sentences should be unaffected by the depth-of-processing effect, especially given that semantic illusions have been hypothesized to involve superficial processing (wason & reich 1979, paape et al. 2020, paape 2021, leivada 2020). however, there was also no evidence in the data that finding meaning in bullshit and illusion sentences correlates with deeper processing, which might be the case if the sentences can be coerced or “repaired” into meaningfulness via additional reasoning steps and/or linguistic operations (dalton 2016, o’connor 2015, wellwood et al. 2018, zhang et al. 2022). it is possible that depth-of-processing effects are simply more subtle for “semi-meaningful” sentences than for clearly sensible or clearly nonsensical sentences. thus, further research with larger participant samples, and ideally with more measurements per participant, is clearly needed. overall, despite the lack of conclusive results, the present work has highlighted the untapped potential of psycholinguistic investigations into bullshit processing. even though it has proven challenging to define bullshit in terms of linguistic features (e.g., cohen 2002), there are some notable tendencies, such as the heavy use of nouns and of abstract rather than concrete words (turpin et al. 2019, buekens & boudry 2015), as well as the heavy use of “genre”-specific jargon (spicer 2020). each of these features may make unique contributions to bullshit processing, as well as to individual differences in bullshit receptivity. applying the entirety of the empirical (psycho)linguistic toolbox to bullshit sentences and investigating connections with other (psycho)linguistic phenomena will undoubtedly lead to many valuable insights in this domain. acknowledgments. the study was funded by the university of potsdam. the author would like to thank the vasishth lab members for helpful comments and suggestions. references de almeida, roberto g. 2018. composing meaning and thinking. in gerhard preyer (ed.), beyond semantics and pragmatics, 201–229. oxford: oxford university press. bainbridge, timothy f, joshua a quinlan, raymond a mar & luke d smillie. 2019. openness/intellect and susceptibility to pseudo–profound bullshit: a replication and extension. european journal of personality 33(1). 72–88. https://doi.org/10.1002/per.2176. barr, dale j., roger levy, christoph scheepers & harry j. tily. 2013. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language 68(3). 255–278. https://doi.org/10.1016/j.jml.2012.11.001. van den broek, p., c. m. bohn-gettler, p. kendeou, s. carlson & m. j. white. 2011. when a proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 198 https://doi.org/10.3765/elm https://www.elm-conference.net/ reader meets a text: the role of standards of coherence in reading comprehension. in m. t. mccrudden, j. magliano & g. schraw (eds.), text relevance and learning from text, 123–139. greenwich, ct: information age. van den broek, paul, robert f lorch, tracy linderholm & mary gustafson. 2001. the effects of readers’ goals on inference generation and memory for texts. memory & cognition 29(8). 1081–1087. https://doi.org/10.3758/bf03206376. buekens, filip & maarten boudry. 2015. the dark side of the loon. explaining the temptations of obscurantism. theoria 81(2). 126–142. https://doi.org/10.1111/theo.12047. čavojová, vladimı́ra, ivan brezina & marek jurkovič. 2020. expanding the bullshit research out of pseudo-transcendental domain. current psychology 41. 827–836. https://doi.org/10.1007/s12144-020-00617-3. christianson, kiel. 2016. when language comprehension goes wrong for the right reasons: goodenough, underspecified, or shallow language processing. quarterly journal of experimental psychology 69(5). 817–828. https://doi.org/10.1080/17470218.2015.1134603. cohen, gerald a. 2002. deeper into bullshit. in lee overton & sarah buss (eds.), contours of agency: essays on themes from harry frankfurt, 321–339. cambridge, ma: mit press. cook, paul & suzanne stevenson. 2010. no sentence is too confusing to ignore. in proceedings of the 2010 workshop on nlp and linguistics: finding the common ground, 61–69. dalton, craig. 2016. bullshit for you; transcendence for me. a commentary on “on the reception and detection of pseudo-profound bullshit”. judgment and decision making 11(1). 121–122. de dios-flores, iria. 2016. more people have presented in conferences than i have. comparative illusions: when ungrammaticality goes unnoticed. in aitor ibarrola-armendariz & jon ortiz de urbina arruabarrena (eds.), on the move: glancing backwards to build a future in english studies, 219–228. bilbao: universidad de deusto, servicio de publicaciones. deyoung, colin g, rachael g grazioplene & jordan b peterson. 2012. from madness to genius: the openness/intellect trait domain as a paradoxical simplex. journal of research in personality 46(1). 63–78. https://doi.org/10.1016/j.jrp.2011.12.003. diamond, cora. 1981. what nonsense might be. philosophy 56(215). 5–22. https://www. jstor.org/stable/3750713. evans, anthony, willem sleegers & žan mlakar. 2020. individual differences in receptivity to scientific bullshit. judgment and decision making 15(3). 401–412. fortuin, egbert. 2014. deconstructing a verbal illusion: the ‘no x is too y to z’ construction and the rhetoric of negation. cognitive linguistics 25(2). 249–292. https://doi.org/10.1515/cog2014-0014. fowler, roger. 1969. on the interpretation of ‘nonsense strings’. journal of linguistics 5(1). 75–83. https://www.jstor.org/stable/4175019. frankfurt, harry. 1986. on bullshit. raritan quarterly review 6(2). 81–100. republished as on bullshit by princeton university press, 2005. franklin, michael s, jonathan smallwood & jonathan w schooler. 2011. catching the mind in flight: using behavioral indices to detect mindless reading in real time. psychonomic bulletin & review 18(5). 992–997. https://doi.org/10.3758/s13423-011-0109-6. gelman, a. & j. hill. 2007. data analysis using regression and multilevel/hierarchical models. proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 199 https://doi.org/10.3765/elm https://www.elm-conference.net/ cambridge, uk: cambridge university press. iacobucci, serena & roberta de cicco. in press. a literature review of bullshit receptivity: perspectives for an informed policy making against misinformation. journal of behavioral economics for policy 6(1). 23–40. kruschke, john. 2014. doing bayesian data analysis: a tutorial with r, jags, and stan. new york: academic press. leivada, evelina. 2020. language processing at its trickiest: grammatical illusions and heuristics of judgment. languages 5(3). 29. https://doi.org/10.3390/languages5030029. lewandowski, daniel, dorota kurowicka & harry joe. 2009. generating random correlation matrices based on vines and extended onion method. journal of multivariate analysis 100(9). 1989–2001. https://doi.org/10.1016/j.jmva.2009.04.008. liddell, torrin m & john k kruschke. 2018. analyzing ordinal data with metric models: what could possibly go wrong? journal of experimental social psychology 79. 328–348. https://doi.org/10.1016/j.jesp.2018.08.009. makowski, dominique, mattan s. ben-shachar & daniel lüdecke. 2019. bayestestr: describing effects and their uncertainty, existence and significance within the bayesian framework. journal of open source software 4(40). 1541. o’connor, ellen. 2015. comparative illusions at the syntax-semantics interface. los angeles, ca: university of southern california dissertation . paape, dario. 2021. the role of incremental and superficial processing in the depth charge illusion: experimental and modeling evidence. 10.31234/osf.io/jp2ma. psyarxiv preprint. psyarxiv.com/jp2ma. paape, dario, shravan vasishth & titus von der malsburg. 2020. quadruplex negatio invertit? the on-line processing of depth charge sentences. journal of semantics 37(4). 509–555. https://doi.org/10.1093/jos/ffaa009. palan, stefan & christian schitter. 2018. prolific.ac – a subject pool for online experiments. journal of behavioral and experimental finance 17. 22–27. https://doi.org/10.1016/j.jbef.2017.12.004. pennycook, gordon, james allan cheyne, nathaniel barr, derek j koehler & jonathan a fugelsang. 2016. it’s still bullshit: reply to dalton (2016). judgment and decision making 11(1). 123–125. pennycook, gordon, james allan cheyne, nathaniel barr, derek j koehler, jonathan a fugelsang et al. 2015. on the reception and detection of pseudo-profound bullshit. judgment and decision making 10(6). 549–563. petrocelli, john v, haley f watson & edward r hirt. 2020. self-regulatory aspects of bullshitting and bullshit detection. social psychology 51(4). 239–253. https://doi.org/10.1027/18649335/a000412. pinheiro, j. c. & d. m. bates. 2000. mixed-effects models in s and s-plus. new york: springer. van prooijen, jan-willem, karen m douglas & clara de inocencio. 2018. connecting the dots: illusory pattern perception predicts belief in conspiracies and the supernatural. european journal of social psychology 48(3). 320–335. https://doi.org/10.1002/ejsp.2331. r core team. 2022. r: a language and environment for statistical computing. r foundation for proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 200 https://doi.org/10.3765/elm https://www.elm-conference.net/ statistical computing vienna, austria. https://www.r-project.org/. reichle, erik d, andrew e reineberg & jonathan w schooler. 2010. eye movements during mindless reading. psychological science 21(9). 1300–1310. https://doi.org/10.1177/0956797610378686. rule, alix & david levine. 2012. international art english: on the rise, and the space, of the art world press release. triple canopy 16. 7–30. sanford, anthony j & patrick sturt. 2002. depth of processing in language comprehension: not noticing the evidence. trends in cognitive sciences 6(9). 382–386. https://doi.org/10.1016/s1364-6613(02)01958-7. schad, daniel j., antje nuthmann & ralf engbert. 2012. your mind wanders weakly, your mind wanders deeply: objective measures reveal mindless reading at different levels. cognition 125(2). 179–194. https://doi.org/10.1016/j.cognition.2012.07.004. sperber, dan. 2010. the guru effect. review of philosophy and psychology 1(4). 583–592. https://doi.org/10.1007/s13164-010-0025-0. spicer, andré. 2020. playing the bullshit game: how empty and misleading communication takes over organizations. organization theory 1(2). 2631787720929704. https://doi.org/10.1177/2631787720929704. stan development team. 2022. stan modeling language users guide and reference manual. https://mc-stan.org. version 2.26.8. turpin, martin harry, alexander walker, mane kara-yakoubian, nina n gabert, jonathan fugelsang & jennifer a stolz. 2019. bullshit makes the art grow profounder. judgment and decision making 14(6). 658–670. walker, alexander, martin harry turpin, jennifer a stolz, jonathan fugelsang & derek koehler. 2019. finding meaning in the clouds: illusory pattern perception predicts receptivity to pseudo-profound bullshit. judgment and decision making 14(2). 109–119. wason, peter c & shuli s reich. 1979. a verbal illusion. quarterly journal of experimental psychology 31(4). 591–597. https://doi.org/10.1080/14640747908400750. wellwood, alexis, roumyana pancheva, valentine hacquard & colin phillips. 2018. the anatomy of a comparative illusion. journal of semantics 35(3). 543–583. https://doi.org/10.1093/jos/ffy014. zhang, yuhan, rachel ryskin & edward gibson. 2022. a noisy-channel approach to depth-charge illusions. preprint available at ssrn: https://ssrn.com/abstract=4130042 http://dx.doi.org/10.2139/ssrn.4130042. proceedings of elm 2: 189-201, 2023 dario paape: investigating the connection between bullshit receptivity and susceptibility to semantic illusions. 201 https://doi.org/10.3765/elm https://www.elm-conference.net/ finding the force: a novel word learning experiment with modals anouk dieuleveut, ailís cournane & valentine hacquard* abstract. this study investigates the semantic and pragmatic challenges of acquiring the force of english modals, which express possibility (e.g., might) and necessity (e.g., must). children seem to struggle with modal force through at least age 4, overaccepting both possibility modals where adults would prefer necessity modals, and necessity modals in possibility situations. these difficulties are typically blamed on pragmatic or conceptual immaturity. in this study, we sidestep these immaturity issues by investigating the challenges of modal learning through a novel word learning experiment with adults, for different ‘flavors’ of modals: epistemic (knowledge-based) versus teleological (goal-based), and comparing novel modals with actual english modals. we find that, when learning possibility modals, adult learners behave as expected: they accept novel modals in necessity situations, both in epistemic and teleological contexts, but less often after they have learned a pragmatically more appropriate necessity modal. however, when learning necessity modals, participants manage to learn the right force (i.e., reject them in possibility situations) for epistemic scenarios only; with teleological scenarios, they accept them in possibility situations. we propose that an overlap in modal flavor explains their behavior, specifically, the competition with an ability interpretation in teleological but not epistemic scenarios, which could also contribute to children’s difficulty with necessity modals reported in the acquisition literature. keywords. modal force; modal flavor; novel word learning experiment; scalar implicatures 1. introduction. english modals express either possibility (e.g., might in (1a)) or necessity (e.g., must in (1b)).1 when and how do children figure out the force of their modals, for instance, that might means ‘possible’, and must, ‘necessary’? the previous acquisition literature shows that children struggle with modal force until at least age 4. they both over-accept possibility modals in contexts where adults prefer necessity modals, and they over-accept necessity modals in possibility situations. these difficulties are typically blamed on pragmatic or conceptual immaturity: their over-acceptance of possibility modals is blamed on difficulty computing implicatures, and failure to realize that necessity modal is often more appropriate in necessity situations (noveck 2001; ozturk and papafragou 2015, a.o.). their over-acceptance of necessity modals in possibility situations is blamed on difficulty reasoning about open possibilities (ozturk and papafragou 2015; see also acredolo and horobin 1987). in both cases, it is assumed that children know the force of the modals, but have difficulty using them. but, could children’s over-acceptance of both possibility and necessity modals instead reflect a lack of knowledge of their underlying force? in this study, we investigate the semantic and pragmatic challenges of modal force acquisition beyond issues of * we thank the members of umd/nyu modality group, s-lab and acquisition lab, especially alexander williams, jeff lidz, as well as floriane dieuleveut for the materials. this research is supported by nsf grant #bcs1551628. authors: anouk dieuleveut, university of maryland (adieulev@umd.edu); ailís cournane, new york university (cournane@nyu.edu) & valentine hacquard, university of maryland (hacquard@umd.edu). 1 further force distinctions can be made: in particular, necessity modals are often split into strong (must) vs. weak (should) necessity (von fintel and iatridou 2008). in this study, we focus on the main contrast between possibility and necessity. proceedings of elm 1: 136-146, 2021 c©2021 anouk dieuleveut, ailı́s cournane and valentine hacquard published by the lsa with permission of the author(s) under a cc by license. 136 https://doi.org/10.3765/elm https://www.elm-conference.net/ conceptual and pragmatic immaturity, by testing how adults learn novel modals, and what situations might be particularly challenging. imagine a child hearing a new modal, sig, in (2). how does she determine whether sig means possible or necessary? (1) a. the keys might be in the drawer. possibility b. the keys must be in the drawer. necessity (2) the keys sig be in the drawer. ? this mapping may be especially challenging for necessity modals, as modals give rise to a classic “subset problem” (xu and tenenbaum 2007; piantadosi et al. 2013; rasin and aravind 2020, a.o.): necessity entails possibility. thus, situations of necessity are also situations of possibility. if sig means possible, but children initially think it means necessary, they should have evidence that their hypothesis is wrong, as sig will be sometimes used in situations where a necessity is logically false (any situations of mere possibility). however, if sig means necessary, but children think it means possible, they may have no counterevidence that their hypothesis is too weak, as sig p will only be used by speakers in situations where a possibility statement, might p, is also logically true.2 the child also needs to figure out how modals are used in conversation, and when the use of one is more appropriate than the other. english has both possibility and necessity modals, which form horn scales (horn, 1972). as such, they can give rise to scalar implicatures (si) (grice 1975; horn 1972): for instance, the use of (1a) can implicate that it is not necessarily the case that the keys are in the drawer (not (must p)). in the gricean tradition, this implicature arises from the assumption that participants in a conversation are trying to be maximally informative: speakers should prefer to use must p if it is relevant. listeners can thus infer that it is not the case that the speaker believes the more informative (logically stronger) sentence (1b): not (must p). to make scalar implicatures, children need to know that their language has dual pairs, i.e., that , similarly to , form a horn scale (horn 1972). however, this is not the case in all languages: ‘variable force’ modals (i.e., modals that are used both in possibility and in necessity situations) have been described for several languages, and analyzed as underlying possibility modals (e.g., nez perce o’qa, deal 2011; gitksan =ima, matthewson 2013; peterson 2010) as well as underlyingly necessity modals (e.g., in st’´at’imcets, rullmann et al. 2008; in washo, bochnak 2015) (see yanovich 2013 for a summary). this means that learners cannot expect their language to have modal duals. and even in a language with duals like english, knowing the force of one modal does not guarantee that the next modal expresses a different force, given that several lexemes can express the same force (e.g., can, might, may). thus, to be able to know that a possibility modal is inappropriate in a necessity situation, the child needs to be aware of the existence (and underlying force) of an alternative necessity modal. now imagine a child who has acquired possibility modals in her language, but has wrongly acquired possibility meanings for her necessity modals, because of the overlap in force (subset problem). this child should thus both accept necessity modals (which she believes express possibility) in possibility situations, and possibility modals in necessity situations (since she lacks a stronger dual). such a mapping error with necessity modals could explain children’s difficulties in previous studies. 2 unless children can use evidence from downward-entailing environments, which reverse patterns of entailment, as suggested for other subset problems (gualmini and schwarz 2009). however, it is unlikely that children can use such a strategy in the case of modals, given their input (see dieuleveut et al. 2019, in prep, for discussion). proceedings of elm 1: 136-146, 2021 anouk dieuleveut, ailı́s cournane and valentine hacquard: finding the force: a novel word learning experiment with modals. 137 https://doi.org/10.3765/elm https://www.elm-conference.net/ we would like to highlight one last potentially complicating factor for force acquisition. in english, as in many other languages, the same modals can express different ‘flavors’, or types, of modality (kratzer 1981, 1991): an epistemic (based on knowledge or evidence, as in (1a-b)), or various root modalities, like teleological (based on goals, and illustrated in (3a-b)), deontic (based on rules), bouletic (based on desires) or ability (based on capacities). the context typically helps determine what flavor is intended. but, from a learner’s perspective it is important to note that these flavors are not mutually exclusive (e.g., it could both be likely and required for my keys to be the drawer). the force challenge could in principle be easier to resolve in some flavors than others. but the overlap in flavor raises a potential further challenge for force acquisition: could a necessity modal intended in one flavor be interpreted as a possibility modal in another flavor? (3) a. to get to the library, we can go down the yellow road. possibility b. to get to the library, we must go down the yellow road. necessity in this study, we investigate the semantic and pragmatic challenges of acquiring the force of modals, independent of conceptual or pragmatic immaturity, by using a novel word learning experiment with adults. given the subset problem, how good are learners at figuring out force? do learners accept novel modals learned in possibility contexts in necessity situations? do they accept novel modals learned in necessity contexts in possibility situations? what is the effect of knowing a scalemate? and finally, do we find differences between flavors, here epistemic vs. teleological? our results show that when learning possibility modals, adult learners behave as expected: they accept novel modals in necessity situations, but less so after they have learned a pragmatically more appropriate necessity modal, both in epistemic and teleological contexts. however, when learning necessity modals, participants manage to learn the right force (rejecting them in possibility situations) only for epistemic scenarios. with teleological scenarios, they do not learn necessity modals, and accept them in possibility situations. we propose that their behavior in teleological scenarios can be explained by an overlap in flavor, specifically, the competition with an ability interpretation. we relate this finding to children’s reported difficulty with necessity modals in the acquisition literature, and argue that overlap in flavor (here, teleological and ability) might be more of a problem than overlap in force (possibility and necessity) for modal force acquisition. 2. study. in this study, we test how adults learn novel modals, by introducing them to a fictional dialect of english. they are taught modals either in possibility contexts (where it is clear that more than one option is open) or necessity contexts (where it is clear than there is just a single option), and then have to judge whether that modal is appropriate in necessity or possibility situations. we investigate how being taught a scale-mate in a second round of learning affects their judgments (e.g., they should be more reluctant to accept a possibility modal in a necessity situation if they’ve previously learned a necessity modal), and whether their judgements vary as a matter of flavor (we tested both teleological and epistemic scenarios, between subjects). 2.1. procedure. participants were introduced to luke, who is learning new words from a (fictional) foreign dialect, kabberton english, from a native-speaker, mary.3 all participants learned 3 novel words in blocks. word 1 was a control (frimp ~ ‘to grab’). words 2 & 3 varied between 3 examples of the experiments (for epistemic and teleological conditions) can be accessed below: kabberton experiment epistemic: http://spellout.net/ibexexps/ad/kab_sg_enep_g_2/experiment.html teleological: http://spellout.net/ibexexps/ad/kab_sg2_tptn_s_1/experiment.html english experiment epistemic: http://spellout.net/ibexexps/ad/kab_mod_enep_s_1/experiment.html teleological: http://spellout.net/ibexexps/ad/kab_mod2_tntp_s_2/experiment.html proceedings of elm 1: 136-146, 2021 anouk dieuleveut, ailı́s cournane and valentine hacquard: finding the force: a novel word learning experiment with modals. 138 https://doi.org/10.3765/elm https://www.elm-conference.net/ subjects, either a novel possibility modal first (learned in possibility situations, tested in necessity situations), or a novel necessity modal first (learned in necessity, tested in possibility), followed by a block with the other modal. for all words, participants first saw mary use the word (training 1: 4 items), then luke (training 2: 4 yes-control, 4 no-control, with feedback), then the test phase (test: 6 trials, 6 yes-control, 6 no-control, no feedback). participants had to judge whether luke’s use of the word was correct, by choosing between ‘yes, that’s correct’ and ‘no, that’s not correct.’ figure 1 illustrates possibility, necessity and impossibility situations by flavor, with test sentences (4).4 figure 2 summarizes the procedure. we ran an identical experiment with english modals as a control, testing might/must (epistemic flavor) and can/must (teleological flavor). the only difference was in the instructions: in the english version, luke was from kabberton, and was learning english with mary. experiments were run on alex drummond’s ibex farm. at the end of the novel word experiment, participants were asked to provide translations. we tested flavor (teleological vs. epistemic) between subjects and force (possibility vs. necessity) within. participants saw either a necessity modal first or a possibility modal. conditions are summarized in (5). details of the instructions are given in the appendix. (4) epistemic: kabberton: ‘the keys sig/gleeb be in the [blue] box.’ english (control): ‘the keys might/must be in the [blue] box.’ teleological: kabberton: ‘we sig/gleeb go down the [blue] road.’ english (control): ‘we can/must go down the [blue] road.’ (5) experiment: kabberton (novel word) vs. english (control) (between subjects) flavor: epistemic vs. teleological (between subjects) force: possibility modal vs. necessity modal (within subjects) order: learnt first vs. second (i.e., knowing a dual) (between subjects) figure 1. visual stimuli and test sentence frames, for epistemic and teleological conditions, by situation type 4 no-controls always corresponded to impossibility situations, regardless of the force of the modal learnt. possibility situation necessity situation impossibility situation ep is te m ic t h e ke ys _ _ b e in t h e b lu e b o x ↑ ↑ ↑ te le o lo g ic al w e _ _ g o d o w n t h e b lu e ro a d ↑ ↑ ↑ proceedings of elm 1: 136-146, 2021 anouk dieuleveut, ailı́s cournane and valentine hacquard: finding the force: a novel word learning experiment with modals. 139 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2. experimental procedure: novel word (kabberton) experiment in the epistemic condition, with block order: control, possibility, necessity. 2.2. expectations. english experiment. participants should reject possibility modals (might/can) in necessity situations if they assume an si, and accept them if they do not (literal interpretation). they should reject necessity modals (must) in possibility situations. kabberton experiment. when learning a novel modal in possibility contexts, participants should accept it in necessity situations, unless they previously learned a stronger dual. when learning a novel modal in necessity contexts, participants should reject it in possibility situations, if they learned it as a necessity modal, but accept it if they learned it as a possibility modal. 2.3. participants. 386 u.s. english participants were recruited on amazon mt (kabberton experiment, n=194 (97 female, age m=37yrs); english experiment: n=192 (97 female, age m=38yrs); after exclusion on controls (3.4%): 373 participants (kabberton: 188, english: 185). 2.4. results. data analyses were conducted using r (r core team, 2013), using the package lme4 (bates et al., 2014a, 2014b). we used binomial linear mixed effects models, built with a maximal random effect structure based on subjects and items as random variables, even though we sometimes had to step back to random intercepts only models when the model failed to converge with the full random effects specification (following barr et al.).5 figure 3 shows proportion of yes responses on test trials for both kabberton and english experiments, depending on block order and flavor, for possibility and necessity modals. test trials for possibility modals correspond to necessity situations (red bars), and to possibility situations (yellow bars) for necessity modals. table 1 reports the mean proportion of yes answers (with standard error) on possibility and necessity situations, depending on order and flavor. analysis. the error rate on controls was very low (kabberton: 3.3%; english: 2.8 %). in the english version of the experiment, we find a relatively low rate of scalar implicatures, especially in the teleological condition (yes answers in necessity: epistemic might: 90.4%; teleological can: 97.9%). participants correctly reject 5 these cases are indicated with ftc (for “failure to converge”) in tables 2 and 3. proceedings of elm 1: 136-146, 2021 anouk dieuleveut, ailı́s cournane and valentine hacquard: finding the force: a novel word learning experiment with modals. 140 https://doi.org/10.3765/elm https://www.elm-conference.net/ necessity modals in possibility situations, for both flavors (epistemic must: 13.0%; teleological must: 19.9%). in the novel word experiment, we find that participants accept novel possibility modals in necessity contexts at high rates for both epistemic (81.5%) and teleological (98.6%) modals, with no difference with the english controls might/can (table 2). however, when they learn necessity modals, we find an unexpected difference between flavors. participants correctly learn a necessity modal in the epistemic condition (they accept it in possibility at only 23.6%, with no significant difference with english might). but they do not seem to learn the necessity modal as necessity in the teleological condition, and accept it 77.2% in possibility situations, with a significant difference with the english control (kabberton vs. english: χ2(1)= 77.9, p <.0001***), suggesting they have learned a possibility modal (table 2). effect of order. for possibility modals, we find a significant effect of order in both experiments: they are accepted less often when learned second, i.e., when subjects know a scale-mate (novel word: epistemic: 46.4%; teleological: 64.3%; english: might: 47.1%; can: 87.2%) (table 3). however, for necessity modals we again find a difference between epistemic and teleological flavors: there is no significant effect for epistemics, and for teleologicals, the effect goes in the opposite direction for the kabberton and the english experiments (decrease for kabberton, increase for english), with a highly significant interaction effect. figure 3. proportion of yes responses for possibility and necessity modals in necessity (red) and possibility (yellow) situations, faceted for force (possibility, necessity), order (1st, 2nd) and flavor (epistemic, teleological) (epistemic: kabberton: n=91 participants *6 observations, english: n=91*6; teleological: kabberton: n=97*6, english: n=96*6). proceedings of elm 1: 136-146, 2021 anouk dieuleveut, ailı́s cournane and valentine hacquard: finding the force: a novel word learning experiment with modals. 141 https://doi.org/10.3765/elm https://www.elm-conference.net/ te le o lo g ic al p 100% (0.0) 98.6% (0.01) 99.7% (0.00) 64.3% (0.07) 99.3% (0.00) 97.9% (0.01) 100.0% (0.00) 87.2% (0.05) n 77.2% (0.06) 99.0% (0.01) 56.9% (0.07) 97.2% (0.01) 19.9% (0.05) 98.6% (0.01) 28.4% (0.06) 98.9% (0.01) table 1. proportion of yes answers (standard error in parentheses) on possibility (p) and necessity (n) situations for possibility and necessity modals, depending on order and flavor, for the kabberton (n=188 participants*6 observations) and the english experiments (n=185 participants*6 observations). bold cells correspond to test cases. accuracy on no-controls (impossibility) was very high, with no difference between groups, so we don’t report here. word learnt 1st word learnt 2nd epistemic possibility ftc with full specification χ 2 (1) = 0.20, p = 0.65 χ 2 (1) = 2.2, p = 0.14 epistemic necessity χ 2 (1) = 3.23, p = 0.072 χ 2 (1) = 0.91, p = 0.34 teleological possibility ftc with full specification χ 2 (1) = 0.46, p = 0.50 χ 2 (1) = 0.55, p = 0.46 teleological necessity χ 2 (1) = 77.9, p <2e-16*** χ 2 (1) = 25.9, p = 4e-07*** table 2. comparison between kabberton and english experiment test conditions, for epistemic and teleological possibility and necessity modals. kabberton experiment english experiment epistemic possibility (tested in necessity) ftc with full specification χ 2 (1) = 21.4, p = 3.7e-06 *** χ 2 (1) = 8.1, p = 0.004 ** epistemic necessity (tested in possibility) χ 2 (1) = 1.37, p = 0.24 (ns) χ 2 (1) = 0.13, p = 0.71 (ns) teleological possibility (tested in necessity) ftc with full specification χ 2 (1) = 10, p = 0.0015 ** χ 2 (1) = 23.5, p = 1.24e-06 *** teleological necessity (tested in possibility) χ 2 (1) = 36, p = 2e-09 *** χ 2 (1) = 38.9, p = 4.34e-10 *** table 3. results of models testing effect of knowing a dual for possibility and necessity epistemic and teleological modals on test conditions for the kabberton and english experiments. kabberton (sig/gleeb) (n=188 * 6) english ({can/might}/must) (n=185 * 6) word learnt 1st word learnt 2nd word learnt 1st word learnt 2nd pos nec pos nec pos nec pos nec ep is te m ic p 99.6% (0.00) 81.5% (0.05) 100% (0.0) 46.4% (0.07) 100.0% (0.00) 90.4% (0.04) 99.6% (0.00) 47.1% (0.07) n 23.6% (0.06) 98.9% (0.01) 48.1% (0.07) 98.9% (0.01) 13.0% (0.05) 98.6% (0.01) 18.9% (0.05) 98.1% (0.01) proceedings of elm 1: 136-146, 2021 anouk dieuleveut, ailı́s cournane and valentine hacquard: finding the force: a novel word learning experiment with modals. 142 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.5. translations. at the end of the novel word experiment, participants were asked to provide translations. table 4 summarizes their answers, by flavor and force. grey cells correspond to accurate translations. we see that possibility modals are overall correctly identified more often than necessity modals. necessity teleological modals are quite often translated with possibility modals (21%), but epistemic modals are not (8%). epistemic necessity modals are quite often translated with be; teleological modals with verbs. table 4. translations for sig/gleeb (kabberton experiment). considered as possibility modals: can, could, may, might, allow, be able; as necessity modals: must, have to, should; other: no answer, i don’t remember; nonsense words. 3. discussion. we find that adult learners accept novel possibility modals in new necessity situations at high rates, for both epistemic and teleological flavors. this suggests that they correctly learn them as possibility modals: they behave as nez-perce speakers, who lack a stronger scalemate in their language that would be more appropriate (deal 2011; see also ozturk & papafragou, 2015). we find a significant effect of having learned a scale-mate, again for both epistemic and teleological flavors: novel possibility modals are accepted less often once learners know that there is another word in the lexicon to describe necessity situations. learners thus make scalar implicatures with novel words. turning to novel necessity modals, we find an (unexpected) difference between epistemic and teleological scenarios. in the epistemic condition, learners correctly learn novel necessity modals: they accept them in possibility situations at only 23.6%, with no significant difference with english must (13.0%). but in the teleological condition, the rate of rejection is much lower: most participants accept them in possibility situations (77%), with a highly significant difference with english can (19.9%), suggesting that they have learnt a possibility modal despite being exposed to the novel modal only in necessity situations in the learning phase. why do participants accept novel necessity modals in teleological possibility situations? and how can we explain the difference between epistemic and teleological conditions? participants’ behavior in the teleological condition might come from differences in perspectives between the learner and the experimenter: participants might not interpret our (intended) teleological necessity situations as such. these scenarios make an ability interpretation salient: the question of whether it is ‘possible or not’ to go down the yellow road might be more salient than whether it is ‘possible or necessary’ to use this road to get to their goal. a competition with an ability interpretation would also explain the ceiling acceptance rate found for english can in those scenarios (98.6% when learned 1st), and the fact that participants do not make scalar implicatures: they do not seem to consider ‘we must go down the yellow road’ as a more relevant sentence to use. potentially reinforcing this, regardless of whether they were learning a possibility or a necessity modal, participants were trained on impossibility situations: this might have indirectly manipulated the qud, and increased the contrast possible/impossible. this result opens up a new possibility for what might make modal force acquisition challenging for children. if children tend to interpret situations where parents intend a teleological necessity as ability, they could lexicalize a possibility meaning for necessity modals. this could explain their p modal n modal be will verb other ep i p 78% 4% 4% 0% 0% 14% n 8% 48% 29% 3% 0% 11% te l p 63% 8% 0% 2% 11% 18% n 21% 53% 0% 3% 11% 14% proceedings of elm 1: 136-146, 2021 anouk dieuleveut, ailı́s cournane and valentine hacquard: finding the force: a novel word learning experiment with modals. 143 https://doi.org/10.3765/elm https://www.elm-conference.net/ difficulties reported in the acquisition literature: both results from comprehension experiments, where children are found to over-accept necessity modals in – intended – possibility situations (e.g., noveck 2001; ozturk and papafragou 2015), and results from corpus studies with younger children, which show that 2-year-olds use necessity modals in situations where adults expect possibility modals (dieuleveut et al. 2019). in epistemic scenarios, the same problem may not arise (at least in our scenarios), as competition with an ability interpretation is less likely.6 4. conclusion. in this study we have addressed semantic and pragmatic factors that may affect modal force learning, using a novel word paradigm with adult english speakers. through studying adult learners, we aim to better understand why child learners struggle with modal force. does their non-adult-like behavior come from conceptual issues, from semantic issues (not having lexicalized the right force, either assuming necessity meanings for might, or possibility meaning for must), or does it come from pragmatics (not making scalar implicatures, i.e. not considering necessity alternatives as relevant when adults would)? our study with adults, who have mature conceptual and pragmatic abilities, suggest that children’s struggles could also be due to problems interpreting flavor correctly, rather than an issue with force per se, nor with implicatures. our results thus highlight the importance of taking into account flavor variability to understand the source of children’s struggles with force reported in the acquisition literature. references acredolo, c., & horobin, k. (1987). development of relational reasoning and avoidance of premature closure. developmental psychology, 23(1), 13. https://doi.org/10.1037/00121649.23.1.13 barr, d. j., levy, r., scheepers, c., & tily, h. j. (2013). random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language, 68(3), 255-278. https://doi.org/10.1016/j.jml.2012.11.001 bates, d., mächler, m., bolker, b., & walker, s. (2014). fitting linear mixed-effects models using lme4. arxiv preprint arxiv:1406.5823. https://doi.org/10.18637/jss.v067.i01 bochnak, m. r. (2015). variable force modality in washo. in proceedings of north-east linguistics society (vol. 45, pp. 105-114). deal, a. r. (2011). modals without scales. language, 87(3), 559-585. https://doi.org/10.1353/lan.2011.0060 dieuleveut, a., van dooren, a., cournane, a., & hacquard, v. (2019). acquiring the force of modals: sig you guess what sig means?. grice, h. p., cole, p., & morgan, j. l. (1975). logic and conversation. 1975, 41-58. https://doi.org/10.1163/9789004368811_003 gualmini, a., & schwarz, b. (2009). solving learnability problems in the acquisition of semantics. journal of semantics, 26(2), 185-215. https://doi.org/10.1093/jos/ffp002 horn, l. r. (1972). on the semantic properties of logical operators. doctoral dissertation, ucla. kratzer, a. (1981). partition and revision: the semantics of counterfactuals. journal of philosophical logic, 10(2), 201-216. https://doi.org/10.1007/bf00248849 6 what makes the ability reading more salient in the teleological case is that a prerequisite for a teleological necessity is that it be circumstantially possible, but this circumstantial possibility is not trivial, so scenarios of teleological necessities necessarily might involve considerations of circumstantial ability. in epistemic necessity scenarios, circumstantial possibility may be more trivially satisfied, and hence not as relevant. proceedings of elm 1: 136-146, 2021 anouk dieuleveut, ailı́s cournane and valentine hacquard: finding the force: a novel word learning experiment with modals. 144 https://doi.org/10.3765/elm https://www.elm-conference.net/ kratzer, a. (1991). modality. semantics: an international handbook of contemporary research, 7, 639-650. https://doi.org/10.1515/9783110842524-004 matthewson, l. (2013). gitksan modals. international journal of american linguistics, 79(3), 349-394. https://doi.org/10.1086/670751 noveck, i. a. (2001). when children are more logical than adults: experimental investigations of scalar implicature. cognition, 78(2), 165-188. https://doi.org/10.1016/s0010-0277(00)001141 ozturk, o., & papafragou, a. (2015). the acquisition of epistemic modality: from semantic meaning to pragmatic interpretation. language learning and development, 11(3), 191-214. https://doi.org/10.1080/15475441.2014.905169 peterson, t. r. g. (2010). epistemic modality and evidentiality in gitksan at the semantics-pragmatics interface. doctoral dissertation, university of british columbia. piantadosi, s. t., tenenbaum, j. b., & goodman, n. d. (2013). modeling the acquisition of quantifier semantics: a case study in function word learnability. under review. rasin, e., & aravind, a. (2020). the nature of the semantic stimulus: the acquisition of every as a case study. natural language semantics, 1-37. https://doi.org/10.1007/s11050-020-091686 rullmann, h., matthewson, l., & davis, h. (2008). modals as distributive indefinites. natural language semantics, 16(4), 317-357. https://doi.org/10.1007/s11050-008-9036-0 searle, j. r. (1969). speech acts: an essay in the philosophy of language (vol. 626). cambridge university press. https://doi.org/10.1017/cbo9781139173438 team, r. c. (2013). r: a language and environment for statistical computing. von fintel, k., & iatridou, s. (2008). how to say ought in foreign: the composition of weak necessity modals. in time and modality (pp. 115-141). springer, dordrecht. https://doi.org/10.1007/978-1-4020-8354-9_6 xu, f., & tenenbaum, j. b. (2007). word learning as bayesian inference. psychological review, 114(2), 245. https://doi.org/10.1037/0033-295x.114.2.245 yanovich, i. (2016). old english *motan, variable-force modality, and the presupposition of inevitable actualization. language, 92(3), 489-521. https://doi.org/10.1353/lan.2016.0045 appendixes instructions. kabberton experiment proceedings of elm 1: 136-146, 2021 anouk dieuleveut, ailı́s cournane and valentine hacquard: finding the force: a novel word learning experiment with modals. 145 https://doi.org/10.3765/elm https://www.elm-conference.net/ english experiment proceedings of elm 1: 136-146, 2021 anouk dieuleveut, ailı́s cournane and valentine hacquard: finding the force: a novel word learning experiment with modals. 146 https://doi.org/10.3765/elm https://www.elm-conference.net/ being tall compared to compared to being tall and being taller jaime castillo-gamboa, alexis wellwood, & deniz rudin* abstract. this paper investigates the semantics of implicit comparatives (alice is tall compared to bob) and its connections to the semantics of explicit comparatives (alice is taller than bob) and sentences with adjectives in plain positive form (alice is tall). we consider evidence from two experiments that tested judgments about these three kinds of sentence, and provide a semantics for implicit comparatives from the perspective of degree semantics. keywords. gradable adjectives; implicit comparatives; degree semantics; experimental semantics 1. introduction. this paper investigates the semantics of implicit comparatives (1), and its connections to the semantics of explicit comparatives (2) and sentences with adjectives in their plain positive form (3):1 (1) alice is tall compared to bob. (2) alice is taller than bob. (3) alice is tall. we first present the results of two experiments probing the empirical profile of these three categories of sentence. we then leverage our results towards the development of a semantics for implicit comparatives that not only predicts the observed judgments, but interacts in the right ways with theories of explicit comparatives and unmodified positive form adjectives. with respect to the relation between implicit and explicit comparatives, we investigate whether and how these two modes of comparison are distinct. three relevant hypotheses are considered: • hypothesis 1: (1) and (2) are symmetrically entailing • hypothesis 2: (1) asymmetrically entails (2) • hypothesis 3: there is no relation between the acceptance of (1) and (2) one could entertain the thought that (1) and (2) are simply different ways of saying the same thing. such a thought would be falsified by evidence against hypothesis 1. evidence for hypothesis 2 or 3 would be useful in fixing on the specific semantic differences between (1) and (2). with respect to the relation between implicit comparatives and positive form adjectives, we investigate whether the evaluation of the plain positive form adjective is closely related to the evaluation of the implicit comparative. specifically, we consider three hypotheses: *we are grateful to george zhiren ye for his help in executing the experiments, and to maribel romero, alexander williams, nurit matuk-blaustein, and reviewers for and attendees of elm for their feedback. authors: jaime castillogamboa, university of southern california (jcastillog@usc.edu), alexis wellwood, university of southern california (wellwood@usc.edu) & deniz rudin, university of southern california (drudin@usc.edu). 1the distinction between implicit and explicit comparatives is explicitly formulated in kennedy (2009). pearson (2010) introduces a further distinction between ‘strong’ implicit comparatives, as in (1), and ‘weak’ implicit comparatives like alice is taller compared to bob. our focus here is on strong implicit comparatives. our ongoing experimental work examines weak implicit comparatives, however we leave discussion of such extensions to future work. proceedings of elm 1: 078-089, 2021 c©2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin published by the lsa with permission of the author(s) under a cc by license. 78 https://doi.org/10.3765/elm https://www.elm-conference.net/ • hypothesis 4: (1) asymmetrically entails (3) • hypothesis 5: the acceptance of (1) is influenced by the acceptance of (3), despite lack of entailment • hypothesis 6: there is no relation between the acceptance of (1) and (3) hypotheses 4 and 5 are different ways of coding the idea that the computation of (1) is importantly related to that of (3), as we might understand from theories on which a phrase like compared to bob modifies a positive form sentence while -er than bob modifies the adjective.2 evidence for hypothesis 6, then, could provide indirect support for linking the semantics of implicit comparatives more closely with that of explicit comparatives. we present our two experiments in §2. both present participants with distributions of lines ranging in height along parameters that should affect the evaluation of plain positive form adjectives. experiment 1 asks about tallness judgments for each line in each distribution using a sentence like (3), and experiment 2 asks about comparative judgments for pairs of lines using sentences like (1) and (2). to preview our results, (i) we find evidence for an asymmetrical entailment from implicit comparatives to explicit comparatives, rejecting hypotheses 1 and 3 and supporting hypothesis 2, and (ii) we find no relation between the acceptance of implicit comparatives and plain positive form adjectives, rejecting hypotheses 4 and 5 and supporting hypothesis 6. buoyed by these results, in §3 we propose an account on which the semantics of implicit comparatives does not involve the computation of the plain positive form adjective, contra one way of understanding ‘context-setting’ views of implicit comparatives (beck et al. 2004, kennedy 2007). importantly, we do not argue that our findings are incompatible with such views; however, we suggest that the difference-based degree semantics that we put forward (building on work by solt 2009, solt & gotzner 2012, zhang & ling to appear) reflects those findings directly. to preview, (1) comes out as true just in case the difference between alice and bob’s height is greater than the contextual standard for a difference-from-bob’s-height. §4 summarizes, and highlights directions of ongoing work. 2. experiments. we tested judgments about implicit comparatives, explicit comparatives and sentences in the positive form. experiment 1 was designed to provide a baseline measure for the evaluation of plain positive form adjectives, and experiment 2 was designed to probe (i) the relationship between the implicit and explicit comparatives (hypotheses 1-3) and (ii) the relationship between the implicit comparative and adjectives in unmodified positive form (hypotheses 4-6). we recruited 60 unique participants on amazon’s mechanical turk who were each compensated $2.00-$2.40 for up to 12 minutes of their time. we restricted participation to turkers located in the united states with a hit approval rate of 99% or greater and a number of approved hits at 1000 or greater. the studies were conducted in accordance with protocols approved by the university of southern california’s institutional review board. 2.1. methodology. we designed a set of 6 distributions of 12 thin rectangles which instantiated the cross of factors clustering (clust, unclust) and scaling (flat, medium, steep) (see 2one example of a relation other than entailment between implicit comparatives and unmodified positive form adjectives (i.e., a pattern supporting hypothesis 5, not hypothesis 4) would be if implicit comparatives pragmatically implicate that the corresponding unmodified positive form sentence is false (sawada 2009). proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 79 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1). we manipulated clustering to probe the impact of the availability of visually salient clusters on tallness judgments (schmidt et al. 2009), and manipulated scaling as a proxy for just how visually salient those clusters are. figure 1: distributions of lines used in experiments 1 and 2 in experiment 1 (n = 20), participants were asked to evaluate (4) against pictures of our distributions with, one at a time, each line in each distribution highlighted in green: (4) the green line is tall. in experiment 2 (n = 40), participants evaluated (5) and (6) (the factor sentence, manipulated between subjects) by appropriate highlights of each possible combination of lines at a distance of 1, 3, and 5 lines apart (the factor difference) in each of our distributions: (5) the blue line is tall compared to the red line. (6) the blue line is taller than the red line. each of these combinations were tested both in a scenario where the blue line was taller than the red line and in one where the red line was taller than the blue line (the factor winner). 2.2. results. our statistical analyses involved logistic mixed effects regressions with random intercepts for subjects. we report p-values based on model comparisons between a maximal model m and a model m′ that subtracts a targeted variable. all analyses were conducted in r using rstudio and the lme4 package (bates et al. 2014). 2.2.1. experiment 1. we looked for signatures of judgments of tallness for individual lines (henceforth, ‘tallness values’), which could be used as predictors for judgments about comparative tallness in experiment 2. potentially surprising given schmidt et al. (2009) was the lack of effect of clustering on tallness values (p = 0.76; means: clust 0.38, unclust 0.39). we did find an effect of scaling (β = 3.16, se = 1.25, χ2(1) = 6.2, p = 0.013), wherein (4) was accepted less as the distribution grew steeper (flat 0.44, medium 0.38, steep 0.33). these two factors did not interact (p = 0.24). the most significant effect in this dataset was that of the position of the green line in the distribution (β = 7.88, se = 0.46, χ2(1) = 1208.9, p < 0.0001) (see figure 2). proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 80 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2: proportion of tallness values in experiment 1 we take from these results that the next experiment should look for (i) an effect of scaling or (ii) some predictive power for the mean tallness values derivable from this dataset in the comparative judgment data. in order to constitute evidence for hypotheses 4 or 5, we would have to observe any such effects selectively impacting the judgments for implicit comparatives. 2.2.2. experiment 2. we looked for differences in the evaluation of implicit and explicit comparatives, and any relation between their evaluation and the patterns observed in experiment 1. we expected that roughly half of the sentences would mark acceptance and half rejection. thus, before conducting our analyses, we transformed raw ‘yes’ and ‘no’ responses into a measure called ‘matching’: a response matched just in case it marked acceptance of (5) or (6) when the blue line was the longer of the two, or it marked rejection and the red line was the longer. we found an effect of winner (β = −0.4, se = 0.16, χ2(1) = 6.44, p = 0.011) corresponding to participants overall providing more matching responses in scenarios where the red line was the winner (mean matching values: blue won 0.91, red won 0.92). this effect was driven by the implicit comparatives, however: we observed an interaction effect between sentence and winner (β = 1.37, se = 0.31, χ2(1) = 19.2, p < 0.0001) showing a large difference in the same direction for implicit comparatives (blue won 0.89, red won 0.94), but a smaller difference in the opposite direction for explicit comparatives (blue won 0.92, red won 0.90) (see figure 3). we also found an effect of difference (β = 0.18, se = 0.07, χ2(1) = 6.79, p = 0.009), such that responses overall matched more as the difference in position between the blue and red lines increased (means: 1-difference 0.91, 30.92, 50.93). here, too, the effect was driven by the implicit comparative, as revealed in an interaction between sentence and difference (β = −0.32, se = 0.14, χ2(1) = 5.61, p = 0.018): matching responses increased for implicit comparatives along with the difference in position (1-difference 0.90, 30.92, 50.94), while responses to the explicit comparatives did not (1-difference 0.91, 30.91, 50.91) (see figure 4). proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 81 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 3: effect of sentence and winner on matching figure 4: effect of sentence and difference on matching what about potential signatures of tallness judgments on the evaluation of implicit and explicit comparatives? crucially, scaling did not play a predictive role here, neither overall (p = 0.58) nor in interaction with sentence (p = 0.73). we didn’t find any effects of tallness values on responses in experiment 2, either. that is, we attempted analyses comparing sentence, winner and the tallness values for those lines marked red and blue on each trial, and found no significant effects. there did not appear to be any predictive power for raw judgments of tallness on the evaluation of implicit comparatives, just as there wasn’t for explicit comparatives. proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 82 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.3. discussion. the effects depicted in figure 3 suggest that any positive difference between the heights of y and x is sufficient to make x is tall compared to y false. this provides evidence for the claim that an implicit comparative entails the corresponding explicit comparative. but it also suggests that not any positive difference between the heights of x and y is sufficient to make x is tall compared to y true. this provides evidence for the claim that an explicit comparative doesn’t entail the corresponding implicit comparative. the effect depicted in figure 4 suggests that, whereas any positive difference between the heights of x and y is enough to make the explicit comparative x is taller than y true, in order for the implicit comparative x is tall compared to y to be true, that difference must meet a certain threshold. taken together, these findings allow us to reject hypotheses 1 and 3 from §1, and lend support to hypothesis 2. the lack of any observed alignment between judgments of tallness and comparative judgments supports the plain intuition that explicit comparatives x is taller than y neither entail nor are entailed by their positive correspondents x is tall, and furthermore suggests the same for implicit comparatives: x is tall compared to y neither entails nor is entailed by x is tall. these findings provide support for hypothesis 6 and against hypotheses 4 and 5 from §1. any positive difference between the heights of x and y is sufficient to make the explicit comparative x is taller than y true, but the likelihood of our participants judging that x is tall compared to y increased as the difference between x and y increased (see figure 4). because this pattern is gradient, not categorical, it suggests the existence of ‘borderline cases’ of implicit comparatives that are not obviously true and not obviously false. 3. analysis. we can now provide a semantics for implicit comparatives that captures the main findings from our experiments. we first introduce some background assumptions in §3.1 about the semantics of explicit comparatives and plain positive form adjectives. in §3.2, we present our preferred account of implicit comparatives, according to which their semantics involves the interaction between the morpheme pos and what we call ‘difference measure functions’, which result from the interaction between compared to-phrases and gradable adjectives. finally, in §3.3, we show that this account correctly predicts support for hypothesis 2 and hypothesis 6, and the existence of borderline cases for implicit comparatives. 3.1. background on degree semantics. we assume a semantics for gradable adjectives according to which they denote measure functions that map objects to degrees (cresswell 1976, stechow 1984, kennedy 1999, 2007).3 for instance, the adjective tall denotes the function tall, which maps objects onto their heights: (7) jtallk = λx.tall(x) on this approach, gradable adjectives combine with degree morphology to produce properties of individuals. this is precisely what happens in the case of explicit comparatives. in (2), tall combines with -er and the than-clause to produce a property that is instantiated by an object just in case its height is greater than bob’s height: 3the main alternative to degree semantics is delineation semantics (klein 1980, doetjes et al. 2011, burnett 2017). from this perspective, gradable adjectives denote functions from objects to truth-values relative to delineations, which represent different ways of classifying objects from a given domain. for discussion of implicit comparatives in the context of delineation semantics, see rooij (2011a, 2011b) and kennedy (2011, 2019). proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 83 https://doi.org/10.3765/elm https://www.elm-conference.net/ (8) jtaller than bobk = λx.tall(x) > tall(b) explicit comparatives are then given truth conditions as in (9): (9) jalice is taller than bobk = 1 iff tall(a) > tall(b) in general, an explicit comparative x is more g than y is true just in case x’s degree of g-ness is greater than y’s degree of g-ness, where g is the measure function denoted by g. the interpretation of the positive form is less straightforward in the context of degree semantics. since, unlike other adjectives, gradable adjectives do not map individuals onto truth-values, they cannot combine directly with the subject of the sentence. to overcome this difficulty, proponents of degree semantic approaches often maintain that, in the positive form, gradable adjectives occur along with an unpronounced morpheme pos, which denotes a function from measure functions to properties of individuals. for instance, in (3), pos combines with tall and returns a property that is instantiated by an object just in case its height is equal to or greater than a certain standard: (10) jtall posk = λx.tall(x) ≥ st the value of st is determined by a standard function s that maps measure functions onto degrees: (11) jposk = λg.λx.g(x) ≥ s(g) the truth conditions of (3) are then as follows: (12) jalice is tall posk = 1 iff tall(a) ≥ s(tall) two features of sentences in the positive form are relevant for our purposes. on the one hand, they are context-sensitive. for instance, (3) might be true relative to a context where the relevant domain includes only alice’s friends from work and false relative to a context where the relevant domain includes only the members of alice’s basketball team. on the other hand, they are vague. consider, for instance, the context involving the members of alice’s basketball team. relative to that context, it seems impossible to draw a sharp line between those individuals who are tall and those who are not. consequently, some individuals are borderline cases of tall.4 there are different ways of accounting for these features. here we follow kennedy (2007) and assume that the explanation for both of them is to be found in the semantics of pos. let’s start with context-sensitivity. on kennedy’s view, pos determines different standard functions relative to different contexts: (13) jposkc = λg.λx.g(x) ≥ sc(g) these standard functions determine different standard degrees, which explains the context-sensitivity of sentences like (3): (14) jtall poskc = λx.tall(x) ≥ sc(tall) (15) jalice is tall poskc = 1 iff tall(a) ≥ sc(tall) 4this is just an intuitive way of highlighting two features of the positive form and is intended to be compatible with theories that treat vagueness as a form of context-sensitivity, as in fara (2000) and raffman (1996). proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 84 https://doi.org/10.3765/elm https://www.elm-conference.net/ the vagueness of the positive form is explained by the way in which standard functions select a standard degree. given a context c, sc determines a standard “in such a way as to ensure that the objects that the positive form is true of ‘stand out’ in [c], relative to the kind of measurement that the adjective encodes” (kennedy 2007:207). if we were to draw a sharp line between tall and non-tall objects, we would violate the ‘standing out’ requirement. 3.2. a difference-based semantics for implicit comparatives. following fults (2005, 2006:§2.3), we assume that compared to-phrases enter the compositional process and combine directly with the gradable adjective. this contrasts with views that treat compared to-phrases as pragmatic context setters, such as those in beck et al. (2004) and kennedy (2007).5 in this section, we introduce a difference-based account of implicit comparatives, building on prior work by solt (2009:§4.5), solt & gotzner (2012), and zhang & ling (to appear). in section 3.3, we show that it accounts for the main results discussed in section 2. according to the difference-based view, an implicit comparative x is g compared to y is true relative to a context c just in case the difference in g-ness between x and y is greater than or equal to a contextually determined standard, which is selected in such a way that the objects that count as g compared to y in c stand out with respect to the difference between their g-ness and y’s g-ness. for instance, (1) [alice is tall compared to bob] is true relative to c just in case the difference in height between alice and bob is greater than or equal to a certain contextually determined standard, which is selected in such a way that the objects that count as tall compared to bob in c stand out with respect to the difference between their height and bob’s height. in order to obtain these truth conditions compositionally, we proceed as follows. first, compared to-phrases denote functions from measure functions to what we’ll call difference measure functions (‘dmfs’, for short). given a measure function g that maps an object onto its degree of g-ness, the dmf of g with respect to x maps an object onto the difference between its degree of g-ness and x’s degree of g-ness. for instance, compared to bob denotes a function that takes a measure function g and returns the dmf of g with respect to bob, i.e., a function that maps an object onto the difference between its degree of g-ness and bob’s degree of g-ness: (16) jcompared to bobk = λg.g−b, where g−b abbreviates λx.g(x)− g(b) on this picture, compared to-phrases are the result of combining the compared to-phrase with noun phrases such as bob. this leaves us with the following lexical entry for compared to: (17) jcompared tok = λx.λg.g−x the next step is to combine compared to bob with tall. by that process, we obtain the dmf in (18), which maps an object onto the difference between its height and bob’s height: (18) jtall compared to bobk = λx.tall−b(x) 5on these views, the function of compared to-phrases is to fix the context in which the embedded sentence is to be evaluated. thus, an implicit comparative x is g compared to y is true relative to a context c just in case x is g is true relative to a different context c∗ that bears a certain relation to c. according to kennedy (2007), c∗ is just like c except that the only individuals in the domain are those being compared, i.e., x and y. on the other hand, beck et al. (2004) suggest that c∗ is a context just like c except that the value that the standard function corresponding to c∗ assigns to the measure function denoted by g is y’s degree of g-ness. proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 85 https://doi.org/10.3765/elm https://www.elm-conference.net/ as with other measure functions, the dmf in (18) combines with pos and returns a property of individuals, which is instantiated by an object just in case the difference between that object’s height and bob’s height is greater than or equal to the standard assigned to the dmf by the standard function in c: (19) jtall compared to bob posk = λx.tall−b(x) ≥ sc(tall−b) finally, the truth-conditions of (1) are given in (20). (20) jalice is tall compared to bob poskc = 1 iff tall−b(a) ≥ sc(tall−b) that is, alice is tall compared to bob relative to a context c just in case the difference in height between alice and bob is greater than or equal to the standard assigned by the standard function in c to the dmf of tall with respect to bob. 3.3. some features of the difference-based view. here we show that the differencebased view can account for three important features of implicit comparatives, which were suggested by the results from our experiments: (i) x is tall compared to y entails but is not entailed by x is taller than y (hypothesis 2), (ii) there is no relation between the acceptance of x is tall compared to y and x is tall (hypothesis 6), and (iii) x is tall compared to y, just like x is tall, has borderline cases. 3.3.1. being tall compared to compared to being taller. recall that, on our semantics, x is taller than y is true just in case tall(x) > tall(y), i.e. just in case tall(x) − tall(y) is greater than 0. now, according to the difference-based view, x is tall compared to y is true just in case tall(x) − tall(y) is greater than or equal to the contextually determined standard for the dmf tall−y. given the ‘standing out’ requirement imposed by the semantics of pos, it is plausible to assume that no context has 0 as the standard of tall−y. thus, tall(x)− tall(y) being greater than or equal to the standard amounts to tall(x)− tall(y) being greater than 0. that is, to x is taller than y being true. therefore, an implicit comparative entails the corresponding explicit comparative. the converse, however, is not true. the fact that tall(x) − tall(y) is greater than 0 doesn’t entail that it is greater than or equal to the contextually determined standard for tall−y. for tall(x) − tall(y) might be a number between 0 and said standard. therefore, an explicit comparative doesn’t entail the corresponding implicit comparative. in sum, the difference-based view predicts hypothesis 2. 3.3.2. being tall compared to compared to being tall. given a context c, x is tall is true just in case x’s height is greater than or equal to c’s standard for tall, whereas x is tall compared to y is true just in case tall(x) − tall(y) is greater than or equal to c’s standard for tall−y. the following cases show that our semantics correctly predicts x is tall compared to y to neither entail nor be entailed by x is tall: • suppose alice is 160 cm tall and bob is 150 cm tall, and consider a context c where the standard for tall is 170 cm and the standard for tall−y is 5 cm. relative to c, the interpretation of (1) is true, but that of (3) is false. • suppose carol is 180 cm tall and bob is 179 cm tall and consider a context c where the standard for tall is 170 cm and the standard for tall−y is 5 cm. relative to c, the interpretation of carol is tall compared to dan is false, but that of carol is tall is true. proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 86 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.3.3. being tall compared to and pos. our experiments did not find a connection between the evaluation of implicit comparatives and that of plain positive form adjectives. however, implicit comparatives are analytically similar to such occurrences of adjectives in that both involve the interaction between pos and measure functions. whereas in x is tall, pos combines with the measure function tall, in x is tall compared to y, it combines with the dmf tall−y. as a result, we expect implicit comparatives to be both vague and context sensitive: • given a context c, sc maps tall−y onto a standard degree in such a way that those who count as tall compared to y stand out in c in virtue of their difference in height with y. as a result, (i) it is not be possible to draw a sharp line between those who count as tall compared to y and those who don’t, and (ii) there are borderline cases of tall compared to y. • the morpheme pos determines different standard functions relative to different contexts. when these functions are applied to tall−y, different degrees are determined as standard degrees. this makes implicit comparatives context-sensitive. the vagueness of implicit comparative was suggested at the end of §2.3. moreover, although our experiments didn’t provide evidence for the context-sensitivity of implicit comparatives, there are cases that give us reason to think that they are in fact context-sensitive. suppose sadie and fido are two golden retrievers. sadie is 80 cm tall and fido is 78 cm tall. relative to a context where the relevant domain includes all sorts of dogs, (21) might be considered false: (21) sadie is tall compared to fido. however, relative to a context where the relevant domain includes only golden retrievers participating in a very prestigious dog show, (21) might be considered true. it is plausible to explain this phenomenon by maintaining that different contexts provide different standards for dmfs. 4. conclusion. in this paper, we have provided a difference-based semantics for implicit comparatives supported by and reflective of the results from two formal judgment experiments. on our approach, implicit comparatives involve some of the same degree-theoretic compositional elements as sentences with unmodified positive form adjectives, i.e., the morpheme pos and an expression denoting a measure function. this helps to explain some of the similarities between the two kinds of sentence, as well as the differences between implicit and explicit comparatives. at the same time, the fact that expressions of the form tall compared to y denote dmfs helps to account for the entailment from x is tall compared to y to x is taller than y. as we noted above, we understand a difference-based account to more directly capture the empirical data uncovered by our experiments than alternative, ‘context-setter’ views. however, we do not understand our data or discussion to rule out such views. in ongoing work, we are testing the evaluation of the same three kinds of sentence as investigated here with different contextual manipulations (including a no-context variant). we are also investigating the semantics of weak implicit comparatives (see note 1) and its interactions with the semantics of strong implicit comparatives. we anticipate that the results of our subsequent studies will shed additional light on the evaluation of competing approaches to the semantics of comparison. proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 87 https://doi.org/10.3765/elm https://www.elm-conference.net/ references bates, douglas, martin maechler, benjamin m. bolker & steven walker. 2014. lme4: linear mixed-effects models using eigen and s4. r package version 1.1-7. http://cran.rproject.org/package=lme4. beck, sigrid, toshiko oda & koji sugisaki. 2004. parametric variation in the semantics of comparison: japanese and english. journal of east asian linguistics 13. 289–344. https://doi.org/10.1007/s10831-004-1289-0. burnett, heather. 2017. gradability in natural language. oxford: oxford university press. https://doi.org/10.1093/acprof:oso/9780198724797.001.0001. cresswell, maxwell john. 1976. the semantics of degree. in barbara partee (ed.), montague grammar, 261–292. new york, ny: academic press. https://doi.org/10.1016/b978-0-12-5458504.50015-7. doetjes, jenny, katerina soucková & camelia constantinescu. 2011. a neo-kleinian approach to comparatives. in ed cormany, satoshi ito & david lutz (eds.), proceedings of semantics and linguistic theory (salt), 19, 124–141. fara, delia graff. 2000. shifting sands: an interest-relative theory of vagueness. philosophical topics 28(1). 45–81. https://doi.org/10.5840/philtopics20002816. fults, scott. 2005. comparison and compositionality. in john alderete, chung hye han & alexei kochetov (eds.), proceedings of the 24th west coast conference on formal linguistics, 146– 154. simon fraser university. fults, scott. 2006. the structure of comparison. college park, md: university of maryland dissertation. kennedy, christopher. 1999. projecting the adjective. new york: routledge. https://doi.org/10.4324/9780203055458. kennedy, christopher. 2007. vagueness and grammar. linguistics and philosophy 30. 1–45. https://doi.org/10.1007/s10988-006-9008-0. kennedy, christopher. 2009. modes of comparison. in malcolm elliott, james kirby, osamu sawada, eleni staraki & suwon yoon (eds.), papers from the 43rd annual meeting of the chicago linguistic society. volume i: the main session, 139–163. chicago: chicago linguistic society. kennedy, christopher. 2011. vagueness and comparison. in paul égré & nathan klinedinst (eds.), vagueness and language use, 73–97. basingstoke: palgrave macmillan. https://doi.org/10.1057/9780230299313. kennedy, christopher. 2019. the sorites paradox in linguistics. in sergi oms & elia zardini (eds.), the sorites paradox, 246–262. cambridge: cambridge university press. https://doi.org/10.1017/9781316683064. klein, ewan. 1980. a semantics for positive and comparative adjectives. linguistics and philosophy 4. 1–45. https://doi.org/10.1007/bf00351812. pearson, hazel. 2010. how to do comparison in a language without degrees: a semantics for the comparative in fijian. in martin prinzhorn, viola schmitt & sarah zobel (eds.), proceedings of sinn und bedeutung, 14, 356–372. raffman, diana. 1996. vagueness and context relativity. philosophical studies 81. 175–192. proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 88 https://doi.org/10.3765/elm https://www.elm-conference.net/ rooij, robert van. 2011a. implicit versus explicit comparatives. in paul égré & nathan klinedinst (eds.), vagueness and language use, 51–72. basingstoke: palgrave macmillan. https://doi.org/10.1057/9780230299313. rooij, robert van. 2011b. vagueness and linguistics. in giuseppina ronzitti (ed.), vagueness: a guide, 123–170. dordrecht: springer. https://doi.org/10.1007/978-94-007-0375-9. sawada, osamu. 2009. pragmatic aspects of implicit comparison: an economy-based approach. journal of pragmatics 41(6). 1079–1103. https://doi.org/10.1016/j.pragma.2008.12.004. schmidt, lauren, noah goodman, david barner & joshua tenenbaum. 2009. how tall is tall? compositionality, statistics, and gradable adjectives. in niels taatgen & hedderik van rijn (eds.), proceedings of cogsci 31, 3151–3156. solt, stephanie. 2009. the semantics of adjectives of quantity. new york, ny: city university of new york dissertation. solt, stephanie & nicole gotzner. 2012. experimenting with degree. in anca chereches (ed.), proceedings of salt 22, 166–187. stechow, arnim von. 1984. comparing semantic theories of comparison. journal of semantics 3(1-2). 1–77. https://doi.org/10.1093/jos/3.1-2.1. zhang, linmin & jia ling. to appear. the semantics of comparatives: a difference-based approach. journal of semantics . proceedings of elm 1: 078-089, 2021 jaime castillo-gamboa, alexis wellwood, and deniz rudin: being tall compared to compared to being tall and being taller. 89 https://doi.org/10.3765/elm https://www.elm-conference.net/ definitely islands? experimental investigation of definite islands anissa neal & brian dillon* abstract. experimental work on islands has used formal acceptability judgment studies to quantify the severity of different island violations. this current study uses this approach to probe the (in-)violability of definite islands, an understudied island, in offline and online measures. we conducted two acceptability judgment studies and find a modest island effect. however, rating distributions appear bimodal across definites and indefinites. we also conducted a self-paced reading experiment, but found no significant effects. overall, offline, definite islands differ from other uniform islands, but online, the results are more complicated. keywords. syntactic islands; definiteness; psycholinguistics 1. introduction. this paper aims to investigate the offline and online processing status of an understudied class of island: definite islands. by definite islands, we mean the apparent ban on extraction from inside definite determiner phrases that was first observed in ross’ dissertation (ross 1967). (1) a. who did irina see a picture of ? b. *who did irina see that/his picture of ? in this work, we will first consider offline judgments to determine how acceptable speakers find these constructions in isolation, and then move to online processing to test whether speakers are sensitive to definite islands in real time. 1.1. definite islands. one interesting feature of definite islands is their somewhat variable status (chomsky 1973). explanations for this gradience vary (chomsky 1973, keller 2000, davies & dubinsky 2003), but it is a shared intuition the example below is intermediate in terms of acceptability (chomsky 1973). (2) ?who did irina see the picture of ? one class of accounts characterizes these as syntactic violations. the dp may be a bounding node that blocks extraction (chomsky 1977, 1973, davies & dubinsky 2003, huang 2018). however, extraction from a dp is possible if certain criteria are met. (3) who did irina write the cruel article about ? for example, davies and dubinsky argue that in (3) the presence of a verb of creation, write, and a semantically related result nominal, article, can override the blocking effect of the dp through abstract noun incorporation. huang (2018) suggests an account that relies on a bound possessor with unvalued phi-features allowing wh-movement from the dp. other accounts take a more semantic *many thanks to maayan keshev, dave kush, ana arregui, and umass’s psycholinguistic and semantics workshops. authors: anissa neal, university of massachusetts, amherst (anneal@umass.edu) & brian dillon, university of massachusetts, amherst (bwdillon@umass.edu). proceedings of elm 1: 237-248, 2021 c©2021 anissa neal and brian dillon published by the lsa with permission of the author(s) under a cc by license. 237 https://doi.org/10.3765/elm https://www.elm-conference.net/ approach, such as the work done by simonenko (2015) on definite dps in austro-bavarian german. simonenko shows that extraction from strong definites creates an uninformative statement by presupposing the content of the possible answers. this provides one potential explanation for their unacceptability, under the assumption that unacceptability can result from an uninformative statement. definite islands are also predicted by more recent-discourse based accounts. the background constituents are islands by goldberg (2013), as the name suggests, posits that backgrounded elements cannot be extracted, and are thus islands. goldberg considers background elements to be constituents that are neither part of the focus domain nor the primary topic. elements that are part of a presupposed clause could then be considered backgrounded. as she and others note, the position of the gap must be within the asserted content of the utterance to be a licit extraction (erteschik-shir 1973), and cannot be presupposed. assuming definite dps are presupposed content, it would be unacceptable to subextract from definite dps under this account. lastly, hofmeister & sag (2010) also investigate referential processing inside the dp. they note that specificity and/or referentiality may consume processing resources, and that filler-gap dependencies are more easily processed when the intervening material is less complex. definiteness may add another layer of difficulty, since a definite dp identifies and situates a specific referent in a discourse model. on this view, the low acceptability for extraction from definite dps is thought of as a reflection of an increased processing toll for filler-gap dependencies in this environment. 1.2. processing islands. much psycholinguistic work has investigated if the parser posits gaps inside island structures (phillips (2006) for a more comprehensive review), and the bulk of it has used a small subset of islands (e.g., relative clause, complex np, wh-islands). in a study very similar to the current one, tollan & heller (2015) found sensitivity to definite islands in an offline task but not in an online task. they investigated two key points: (i) how sensitive is the parser to definiteness in processing filler-gap dependencies, and (ii) if, and how, the type of wh-phrase influences the parsing of the dependency. their first experiment was an online, self-paced reading task using the filled-gap effect paradigm. the filled-gap effect refers to the finding that readers have difficulty when another word is already in a position where they expect to see a gap (stowe 1986). this paradigm can be used to determine where readers are actively positing possible gap sites as they parse a sentence. examples of their stimuli are below. they also include a manipulation of d-linked vs. non d-linked fillers. (4) a. which singer did lizzie see [a/the movie about elvis presley] with ? b. who did lizzie see [a/the movie about elvis presley] with ? c. did lizzie see [a/the movie about elvis presley] with kate bush? participants were given a preceding context that made the question felicitous, and were then presented with the sentences word-by-word. yes/no questions were included as a baseline condition to measure the size of the filled-gap effect. the authors predicted no effect of definiteness in the first gap position at “a/the movie,” and compare only the d-linked and non d-linked items. they find a significant slowdown in this region for the d-linked phrase, suggesting participants had greater difficulty and hence a greater filled-gap effect, with which than who. in the second gap position, they find no definiteness effect in the yes/no questions, and focus on the filled-gap proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 238 https://doi.org/10.3765/elm https://www.elm-conference.net/ sentences. they do find a main effect of wh-type with which-nps having slower reading times. however, there is no effect of definiteness. the authors also ran an offline, question completion study. (5) a. which singer did lizzie see a/the movie about b. who did lizzie see a/the movie about c. did lizzie see a/the movie about participants were given the same short, preceding contexts as experiment 1 and then asked to complete the sentence prompts. ending the sentence by adding a question mark was a possible response. the authors reasoned that if the participants find a gap inside an island acceptable then they should add a question mark after about in all cases. however, if this is an unacceptable gap site, they may continue the sentence to place the gap outside the dp (i.e., who did lizzie see the movie about elvis presley with ?). yes/no questions were once again used as a control. the offline results showed evidence for a definiteness effect: participants placed gaps inside indefinite nps more than definite nps. thus the authors found mixed evidence of sensitivity to definiteness in processing. furthermore, they use the same methodology used in this study, an online self-paced reading study paired with an offline task. however, one of the main differences between this study and the one done in this paper, aside from the d-linking manipulation, is different baselines. in the tollan and heller study, they use yes/no questions as their control condition for both the offline and online studies. while yes/no questions and wh-questions share certain similarities, there are a variety of syntactic and semantic differences that could disrupt the interpretation of the results. the controls in the present study manipulate only the presence of a filler, given in (7). the tollan and heller results also raise the possibility that the parser will entertain gaps inside definite islands in real-time processing, in apparent violation of a grammatical constraint. this would be a surprising conclusion in light of the broader literature, and so bears further scrutiny. therefore, it is crucial to see if the tollan and heller results replicate, and if there is actually a distinction between the offline and online results. 2. experiment 1: acceptability judgment. as seen in (1), there is a range of judgments associated with definite islands. the purpose of experiments 1a and 1b is to determine if naive informants share the intuitions reported in the literature. that is, just how acceptable do speakers consider extractions from the-dps to be. to investigate this, we used the factorial paradigm for islands developed by sprouse et al. (2016). this methodology has been successfully applied to a variety of different islands (e.g relative clause, complex np, subject, adjunct, and wh-islands) in a variety of languages. the factorial design is useful for several reasons. it allows for a quantitative measure of an “island effect.” as sprouse et al. (2016) notes in his overview of this design, there many extrasyntactic factors that could go into the decreased acceptability observed in islands. for example, long distance dependencies are considered less acceptable than short distance ones, such as in the example below, where the filler is much further in the second case than in the first. (6) a. indefinite, matrix: tara knows who found a photo of yelena. b. indefinite, embedded: tara knows who fiona found a photo of . proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 239 https://doi.org/10.3765/elm https://www.elm-conference.net/ c. definite, matrix: tara knows who found the photo of yelena. d. definite, embedded: tara knows who fiona found the photo of . furthermore, the complexity of the island itself could impact the acceptability. different types of islands could introduce a certain amount of complexity related to structure, meaning, or both. sentences containing islands may thus, by virtue of the island structure alone, have a lower acceptability. with the factorial design, extra-syntactic factors like dependency length and complexity are independently estimated and accounted for. the effect of length can be measured by subtracting (6d) from (6c). the effect of the island structure can be measured by subtracting (6a) from (6c). the total effect is measured by subtracting (6b) from (6c). the length and structure effects are then subtracted from the total effect, and the remainder is the size of the “island effect.” the island effect is quantified as this difference-in-differences (dd) score, with larger values indicating more severe island penalties. 2.1. experiment 1a. embedded judgment study. participants. 41 native american english participants were recruited for an online acceptability judgment study run on ibexfarm (drummond 2020) through the prolific academic platform and paid $4 each. materials. the study was a 2x2 within-subjects factorial design with factors of distance: long, short and definiteness: indefinite, definite (7) a. the journalist guessed who promoted a/the ridiculous photo of madonna. b. the journalist guessed who charlie promoted a/the ridiculous photo of . there were a total of 24 experimental items and 48 filler items, designed it to have approximately equal numbers of acceptable and unacceptable items. of the filler items, 33 of the 48 were ungrammatical, and about half of the ungrammatical fillers shared similar characteristics to the experimental items. some of the fillers were taken from the sprouse et al. (2016) experiment. procedure. participants were given instructions on how to rate acceptability and tested on three practice sentences. each sentence was presented in full, and the participant was asked to give the provided sentence a rating on a scale from 1 to 7. 1 was the most unacceptable, and 7 was the most acceptable. the experimental lists were constructed in latin square fashion; each participant saw only one experimental token from the above paradigm for each item. they were encouraged to use the full range of the scale in the instructions. participants could complete the experiment at their own pace on ibexfarm; it was expected to take 20 minutes to complete, and the average completion time was 15 minutes. analysis. the participants read a series of questions to ensure they were properly following instructions. if the accuracy of their response was less than or equal to 60%, the participant was excluded. a total of 39 out of 41 participants were analyzed. the data were z-transformed before analysis to account for different scale usages across participant. the factors were sum coded (definiteness: definite = -0.5, indefinite = 0.5; distance: short = -0.5, long = 0.5) and a mixed-effects proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 240 https://doi.org/10.3765/elm https://www.elm-conference.net/ experiment 1a experiment 1b estimate se t p estimate se t p distance -0.04 0.06 -0.73 0.001 0.06 0.07 8.99 <0.001 definiteness 0.79 0.07 11.08 0.46 0.002 0.06 0.028 0.98 definiteness:distance 0.22 0.11 1.92 0.055 0.25 0.11 0.003 0.028 table 1: experiment 1 linear regression results linear regression model using the lmer test of the lme4 package in r (bates et al. 2015, r core team 2020) was fit to the z-scored data. following matuschek et al. (2017), random slopes were removed until convergence 2.1.1. results. figure (1a) presents the interaction plot for the z-scored means. the linear model found a significant main effect of distance and a marginal interaction (table 1). overall, the long filler-gap conditions were less acceptable, but the length penalty was larger for the definite conditions. 2.1.2. experiment 1a discussion. we observed a large effect of distance on acceptability. sub-extraction from inside the dp was much worse than short distance extraction across both types of dps, definite and indefinites. the low ratings for long distance extraction are likely related to the fact that people prefer shorter filler-gap dependencies. previous work on cross-clausal extraction (e.g., mcelree et al. (2003)) indicates that the longer the distance between the filler and the gap, the more processing difficulties arise. the marginal interaction is interesting as it hints at the possibility of an island effect, but it is inconclusive. sprouse et al. (2016) notes that the factorial paradigm produces three ways in which to observer an island effect: the presence of a significant interaction, visual absence of parallel lines on the interaction plot, and a difference-in-differences score that is greater than 0. the results of experiment 1a satisfy all but the first. interestingly, sprouse et al. notes that the magnitude of a dd score is a concern for syntactic theory. while these two of these three points would appear to suggest an island effect in definite dps, the numerically small dd leaves this question open. (a) interaction plot for 1a (b) interaction plot for 1b figure 1: experiment 1 interaction plots proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 241 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.2. experiment 1b: direct judgment study. the second experiment followed the same methodology and design as the first experiment but tested direct wh-questions instead of embedded indirect wh-questions. the reason for this was to test pragmatic-semantic accounts that tie this effect straightforwardly to the semantics of direct questions simonenko (2015). participants. a total of 43 native american english participants were recruited for an online acceptability judgment study run on ibexfarm through the prolific academic platform. as in the first experiment, they were paid $4. filters were put in place to ensure that participants that had completed the previous study could not participate in this study. materials. the stimuli was based off those of experiment 1a, but they were adjusted to be direct questions. the matrix clause was removed, and the embedded verb, which contains the dp, became the main verb. (8) a. who published a/the horrible article about gina? b. who did olivia publish a/the horrible article about ? as in experiment 1a, there were a total of 24 experimental items and 48 filler items. procedure. the procedure was the same as in experiment 1a. analysis. the same exclusion criteria from experiment 1a were also applied. a total of 40 out of 43 participants were analyzed. one item was removed from analysis due to coding error making a total of 23 experimental items. analysis was identical to experiment 1a. 2.2.1. results. table 1b presents interaction plot, and table 1 the results of the linear model. we observed a significant effect of distance and a significant interaction of definiteness and distance, visualized in figure 1b. 2.2.2. experiment 1b discussion. experiment 1b differs from experiment 1a in that there was a significant interaction between distance and definiteness. that is, there appears to be a super-additive effect in the sense that the low acceptability of (8b) cannot be explained through the individual effects of distance or definiteness alone. again, we observed a large effect of distance. we are, however, hesitant to conclude that these results show evidence of an island effect in direct questions, but not in embedded. first, a direct comparison between the two is difficult due to a lack of power. second, while there is a significant interaction in direct questions, one of the key benefits of the sprouse paradigm is a measure of an islands magnitude. the dd score for both experiments are of a lower magnitude than other observed islands (see figure 2), and in fact, are somewhat similar to each other. 2.3. experiment 1 discussion. the dd scores of both experiments were very close, 0.22 and 0.25, respectively. these dd scores are on the lower end of dd scores observed for island constructions when compared to other islands in the literature. while kush et al. (2019) states that there is no true threshold for dd scores to be representative of a real island effect, the range from previous studies fall within 0.75 to 1.25. figure 2 is a summary of dd scores across previous experiments following the same factorial design. results of the current study are presented in red. the dashed-grey line indicates that there was a significant interaction of distance and definiteness proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 242 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2: dd scores across studies. figure 3: participant rating distribution for experiment 1. red indicates definites, and blue indefinites. for the islands above and to the right of the line. the languages used in the studies were hebrew, norwegian, english, and italian, respectively. as seen in figure 2, definite islands maintains a borderline placement compared to other islands. the dd scores for both conditions are on the lower side, patterning more with nonislands in english and other languages. for example, the rc adjuncts from sprouse et al. (2016) have a dd score of 0.01. sprouse and colleague concluded that the apparent unacceptability of rc adjuncts can be explained through an effect of distance, rather than a true “island effect.” while the definite islands produced a larger dd than did rc adjuncts, it is still appears quite small compared to the more robust islands with larger dd scores. this raises the question as to whether the definite islands are akin to other islands, or if the effect could be explained as compounding effects of distance and definiteness. while the results of experiment 1 were instrumental in developing a further understanding of how acceptable these islands are offline, it does not present entirely proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 243 https://doi.org/10.3765/elm https://www.elm-conference.net/ conclusive evidence. experiment 2 investigates if speakers will posit gaps inside definite islands during online processing. given the results of experiment 1, it is unclear what behavior we might observe in definite islands in terms of active gap positing. while previous evidence suggests that gap-filling does not occur inside islands, the nebulous judgments from experiments 1 suggest that these are not particularly strong island environments. furthermore, the results from tollan & heller (2015) suggest that there may be a offline-online asymmetry for definite islands, as they found evidence of a definiteness effect offline but not online. a look at the z-score rating distribution also suggests something interesting. figure 3 is a density plot that shows the z-scored ratings of all the experimental items across the experiments. there appears to be a bimodal distribution for both experiments in the case of long distance extraction. since this figure is based off the z-scored ratings, zero here indicates deviation from the mean for each group. therefore, the bimodality suggests that there are a fair amount of ratings that are higher than average and lower than average for the cases of long distance extraction. for the cases of short distance extraction, they appear to be consistently higher than average. these two patterns are observed for both the definites and the indefinites. this bimodality could be caused by several things. first, perhaps particular items in the experiment could be more acceptable than others. if item-wise variability was driving this bimodal pattern, then we would expect the difference-of-differences score by item in e1a to be predictive of their ratings in e1b, since they share lexical material. to check this, a correlation between the dd scores of each item for the two experiments was done. the resulting pearson correlation was 0.22, and regressing the dd values onto each other was not significant. while this could just reflect low power, to some degree it suggests that the bimodality is not due to item-specific factors. if the item bias was present, it should still show up despite this change. the subjects themselves could also be bimodal. that is, there could be some participants who consistently rate the island conditions above the average and those that do not. to investigate this possibility, a histogram of each participant’s z-scored rating for each condition was plotted. if subjects are patterning bimodally, then the histograms (figure 4) should match the density plot. this does not appear to be the case. (a) 1a: top definite; bottom indefinite (b) 1b: top definite; bottom indefinite figure 4: subject means for long distance extraction proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 244 https://doi.org/10.3765/elm https://www.elm-conference.net/ another key aspect is that bimodality is found in both the definite and indefinite dps. speakers are also not consistently rating the indefinites at average or higher in the long distance extraction cases. while some of this could be due to the fact that, as seen in both experiments 1a and 1b, distance does seem to greatly impact acceptability, it is still interesting to observe this in what are considered grammatical sentences. the fact that neither itemnor subject-wise variability appears to be the source, this raises several questions about the true nature of the bimodality. 3. experiment 2: self-paced reading. experiment 2 is an online measure attempting to address how the parser handles the-dp islands in real time. it is a self-paced reading experiment that uses a filled-gap effect paradigm, discussed in section 1.2. under this design, if participants are actively positing gaps inside the definite islands we should expect to see a slowdown in reading times if a dp-internal gap position is filled. an example stimulus is below. (9) a. the journalist guessed who charlie promoted a a/the ridiculously scandalous photo of b einstein and max planck to . b. the journalist guessed that charlie promoted a/the ridiculously scandalous photo of einstein and max planck to the scientific magazine. there are two possible gap sites for each item. the first, labeled a above, comes directly after the embedded verb, “promoted” (e.g., the journalist guessed who charlie promoted). the second, labeled b, is inside the dp, after “of” (e.g., the journalist guessed who charlie promoted a ridiculously scandalous photo of ). gap a is a baseline to determine is participants are actively positing gaps in general; we predict no effect of definiteness here. gap b investigates whether participants are positing gaps inside the dp; there should be an observed definiteness effect here if definite islands pattern like other islands. participants. 45 american english speakers all recruited through the prolific platform. they were paid $5 for their participation. materials. the stimuli, (9), for the self-paced reading experiment are identical to those used in experiment 1 with a few adjustments to create a filled-gap effect. the name inside the dp was extended with a coordination to allow for spillover, and a continuation was added to provide a grammatical gap site. this resulted in a 2x2 within-subjects design with factors of definiteness: indefinite and definite and phrase type: who and that. there were a total of 24 items, distributed in four latin-squared lists. there were a total of 40 filler items, not including 4 practice items to get participants accustomed to the task. a quarter of the fillers mirrored the experimental items, and the remaining filler items were unrelated. procedure. the experiment was run on ibexfarm using a word-by-word self-paced reading paradigm. participants were given detailed instructions on how to complete the experiment and asked a series of comprehension questions to ensure they were paying attention. after four practice items, the experiment began. participants were presented with a “+” for 1500 ms after which the screen changed to become the first word of the sentence. they then had to press the spacebar to move throughout. once they had completed the entire sentence, they were asked a yes/no question and used the keyboard to answer. the instructions and practice items emphasized the importance proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 245 https://doi.org/10.3765/elm https://www.elm-conference.net/ of reading at a natural pace and trying to understand the sentences in an attempt to ensure participants were paying attention. after the experiment ended, participants were asked to rate a single set of sentences similar to the experimental ones with the binary options of “good” or “bad”. they were taken to a set of post-experiment screening questions, and the experiment concluded. analysis. the data were pre-processed to remove participants who failed to maintain an accuracy of 60% or above on the instruction questions as in experiment 1. participants were also removed if in a post-experiment question their responses to an open-ended question indicated they might be a bot. we also screened for compliance with the experimental instructions to ensure participants were reading the sentences in a word-by-word manner, as opposed to simply ‘clicking through’ at a fixed pace. to do this, two regression models were fit to each participant’s log-transformed rt data. one had only an intercept (the null model), and the other model used region as a predictor. if participants are varying the speed at which they move through the sentence, the the region model should perform better. these two models were fit for each participant and a likelihood ratio test was performed. if we were not able to reject the null model at α = 0.2, then the participant was removed from further analysis. 3.1. results. for both gap sites, a and b, there were no significant effects. the spillover region also did not find any significant results. figure 5: experiment 2 reading times (ms) 1st gap 2nd gap critical spillover critical spillover β p β p β p β p phrase type -0.04 0.08 -0.03 0.19 -0.013 0.58 -0.003 0.86 definiteness -0.01 0.53 -0.02 0.45 0.003 0.87 -0.02 0.33 phrase:definiteness -0.05 0.25 0.03 0.47 0.014 0.73 -0.02 0.53 3.2. experiment 2 discussion. no reliable filled-gap effect was observed at either position. in the first gap, there was a marginal main effect of phrase type (p=0.08). in the second, we fail to see any significant effects. the initial gap was meant as a baseline to ensure that we could proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 246 https://doi.org/10.3765/elm https://www.elm-conference.net/ measure a filled-gap effect with our materials. figure 5 suggests there is a numerical trend towards a filled-gap effect, though it is not reliable in our analysis. this could reflect low statistical power. turning to the gap site inside the dp, gap b, we saw no clear differences in the mean reading times across any of the four conditions. participants, it would appear, do not seem to be positing gaps inside definite or indefinite dps. this observation raises several questions. first, is whether a filled-gap effect is at all present inside a dp; the current results yield no evidence for this. second, is whether this filled-gap effect would be modulated by definiteness. this, once again, was also not found in the current study. however, the early marginal effect in the initial gap position could indicate that our failure to find an effect is due to a lack of power. increasing the power could further clarify both of the above questions. if the reading times remain equal across all four conditions at the dp-internal gap site, this would be evidence to support a claim that participants are not attempting to posit gaps inside dps. in sum, this present study finds no significant evidence of active gap filling inside dps, definite or indefinite. this is somewhat in line with the findings of tollan & heller (2015); they find that there is no effect of definiteness between who and which np. they do, however, observe a filledgap effect inside the which np condition which is not modulated by definiteness. it is also unclear if the extraction cases (i.e., who-phrases and which np-phrases) differed significantly from the non-extraction cases (i.e., yes-no questions). 4. conclusions and future work. definite islands present varied behavior across offline and online experiments. the offline results suggest a weak island effect that is smaller than other observed islands in the experimental syntax literature. it also reveals a bimodal distribution of judgment where participants show varied ratings for both definite and indefinite dp extractions. this bimodality does not appear to be driven by itemor subject-wise variability. in a reading time study, we failed to find evidence for active gap filling inside the dp, making it difficult to test whether definiteness restricts gap filling inside a dp. taken together, the picture that this work presents is that definite islands, and dps in general, consist of greater nuance than originally thought at both an offline and online level. references bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67(1). 1–48. 10.18637/jss.v067.i01. chomsky, noam. 1973. conditions on transformations. in s anderson & p kiparksy (eds.), a festschrift for morris halle, 232–286. new york: holt, rinehart & winston. chomsky, noam. 1977. on wh-movement. in p culicover, t wasow & a akmajian (eds.), formal syntax, 71–132. new york:: academic press. davies, william d. & stanley dubinsky. 2003. on extraction from nps. natural language & linguistic theory 21(1). 1–37. 10.1023/a:1021891610437. https://doi.org/10. 1023/a:1021891610437. drummond, alex. 2020. ibex farm. https://spellout.net/ibexfarm/. erteschik-shir, nomi. 1973. discourse constraints on dative movement. syntax and semantics 12. proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 247 https://doi.org/10.3765/elm https://www.elm-conference.net/ goldberg, adele e. 2013. backgrounded constituents cannot be “extracted”. in jon sprouse & norbert hornstein (eds.), experimental syntax and island effects, 221–238. cambridge: cambridge university press. 10.1017/cbo9781139035309.012. https: //www.cambridge.org/core/product/identifier/9781139035309% 23c00870-943/type/book-part. hofmeister, philip & ivan a. sag. 2010. cognitive constraints and island effects. language 86(2). 366–415. 10.1353/lan.0.0223. http://muse.jhu.edu/content/crossref/ journals/language/v086/86.2.hofmeister.html. huang, nick. 2018. the bound possessor effect: a new argument for the phasehood of definite dps [manuscript]. in north east linguistic society 48, . keller, frank. 2000. experimental and computational aspects of degrees of grammaticality: university of edinburgh phd. kush, dave, terje lohndal & jon sprouse. 2019. on the island sensitivity of topicalization in norwegian: an experimental investigation. language 95(3). 393–420. 10.1353/lan.2019.0051. https://muse.jhu.edu/article/733277. matuschek, hannes, reinhold kliegl, shravan vasishth, harald baayen & douglas bates. 2017. balancing type i error and power in linear mixed models. journal of memory and language 94. 305 – 315. https://doi.org/10.1016/j.jml.2017.01.001. http://www. sciencedirect.com/science/article/pii/s0749596x17300013. mcelree, brian, stephani foraker & lisbeth dyer. 2003. memory structures that subserve sentence comprehension. journal of memory and language 48(1). 67–91. isbn: 0749-596x publisher: elsevier. phillips, colin. 2006. the real-time status of island phenomena. language 82(4). 795–823. 10.1353/lan.2006.0217. http://muse.jhu.edu/content/crossref/journals/ language/v082/82.4phillips.pdf. r core team. 2020. r: a language and environment for statistical computing. r foundation for statistical computing vienna, austria. https://www.r-project.org/. ross, john robert. 1967. constraints on variables in syntax: mit phd. https://eric.ed. gov/?id=ed016965. simonenko, alexandra. 2015. semantics of dp islands: the case of questions. journal of semantics 33. ffv011. 10.1093/jos/ffv011. sprouse, jon, ivano caponigro, ciro greco & carlo cecchetto. 2016. experimental syntax and the variation of island effects in english and italian. natural language & linguistic theory 34(1). 307–344. 10.1007/s11049-015-9286-8. http://link.springer.com/10. 1007/s11049-015-9286-8. stowe, laurie a. 1986. parsing wh-constructions: evidence for on-line gap location. language and cognitive processes 1(3). 227–245. 10.1080/01690968608407062. https://doi.org/10.1080/01690968608407062. publisher: routledge eprint: https://doi.org/10.1080/01690968608407062. tollan, rebecca & daphna heller. 2015. elvis presley on an island: wh dependency formation inside complex np objects. in north east linguistic society 46, https://www. researchgate.net/publication/303306297. proceedings of elm 1: 237-248, 2021 anissa neal and brian dillon: definitely islands? experimental investigation of definite islands. 248 https://doi.org/10.3765/elm https://www.elm-conference.net/ reciprocal predicates: a prototype model imke kruitwagen, yoad winter & james hampton* abstract. many languages have verbal stems like hug and marry whose intransitive realization is interpreted as reciprocal. previous semantic analyses of such reciprocal intransitives rely on the assumption of symmetric participation. thus, ‘sam and julia hugged’ is assumed to entail both ‘sam hugged julia’ and ‘julia hugged sam’. in this paper we report experimental results that go against this assumption. it is shown that although symmetric participation is likely to be preferred by speakers, it is not a necessary condition for accepting sentences with reciprocal verbs. to analyze the reciprocal alternation, we propose that symmetric participation is a typical feature connecting the meanings of reciprocal and binary forms. this accounts for the optionality as well as to the preference of this feature. further, our results show that agent intentionality often boosts the acceptability of sentences with reciprocal verbs. accordingly, we propose that intentionality is another typical semantic feature of such verbs, separate from symmetric participation. keywords. reciprocal predicates; verb meaning; prototype theory; typicality; experimental semantics; lexical semantics. 1. introduction. reciprocal verbs alternate between a unary and a binary form: (1) julia and sam hugged. (2) sam hugged julia. in the intransitive sentence (1) the verb hug takes a plural agent, whereas (2) distinguishes two thematic roles. sentences like (1) and (2) have a clear semantic relation with each other. intuitively, sentence (1) reports a collective activity where the participants perform acts of the sort reported by (2). however, the specifications of this semantic relation are not easy to pinpoint. most previous works have analyzed the semantics of alternations as in (1)-(2) under the assumption that reciprocal events as reported by (1) are obligatorily symmetric. thus, these works rely on the introspective judgement that (1) logically entails (2). in contrast to these truth-conditional approaches, we propose that sentences like (1) support symmetric events as part of their most typical interpretation. thus, (2) is true in the most typical events reported by (1) but it is not logically entailed by (1). this typicality approach is supported by experimental results on filmed scenarios where speakers accept reciprocal sentences (1) but reject sentences like (2). we propose that the typicality approach embodies a more adequate analysis of reciprocal verbs, which is also relevant for other alternations such as active-passive alternations or active-causative alternations. * for their dedicated work on the materials for the experiment, we are thankful to hilbert dijkstra, eva kröse, anna peeters and lianne zandstra. this work was funded by the european research council (erc) under the european union’s horizon 2020 research and innovation programme (grant agreement no 742204). authors: imke kruitwagen, utrecht university (i.kruitwagen@uu.nl), yoad winter, utrecht university (y.winter@uu.nl) & james hampton, city, university london (j.a.hampton@city.ac.uk). proceedings of elm 1: 197-203, 2021 c©2021 imke kruitwagen, yoad winter and james hampton published by the lsa with permission of the author(s) under a cc by license. 197 https://doi.org/10.3765/elm https://www.elm-conference.net/ previous works followed different theoretical lines to analyze the reciprocal alternation. gleitman (1965) and lakoff & peters (1966) use transformational rules that connect the different verbal forms in a reciprocal alternation. dowty (1991) analyzes the alternation by means of proto-roles, where the reciprocal form contains multiple proto-roles for agents, who act on each other according to the proto-roles for the binary form. carlson (1998) on the other hand argues that the unary form comes with only one agent role, so different entities can carry the agent role together. dimitriadis (2008) and siloni (2008) strengthen carlson’s account, claiming that unary reciprocals necessarily denote an irreducible symmetric event. gleitman et al. (1996) focus on the relation between reciprocity and symmetry of binary predicates. they point out that many predicates like hug, kiss, and fight do not show symmetry in their binary form. however, like the other aforementioned works, gleitman et al. also assume that reciprocal entries uniformly require symmetric participation of the agents in the sense that is exemplified above for (1). thus, while there are many technical and theoretical differences between the different approaches to lexical reciprocity, they all share the prediction about a logical entailment between sentences like (1) and (2): any speaker who accepts (1) as true is also expected to accept (2) as true. in short: (3) julia and sam hugged → sam hugged julia and julia hugged sam. we refer to this assumption on reciprocal verbs like hug as symmetric participation. intuitively, it is not clear that all reciprocal verbs indeed show symmetric participation in this way. consider sentence (4) below, in a situation where a car collided with a stationary truck: (4) the car and the truck collided. postulating symmetric participation for (4) would expect it to be judged false in this situation. however, when we informally consulted speakers about (4), the majority of them said they would accept (4) in the given scenario. to account for such judgements, our proposal challenges the assumption that sentences with unary reciprocals logically entail symmetric participation of subject entities. instead of treating symmetric participation as necessary for the truth of sentences with reciprocal verbs, we hypothesize that it is only a preferred condition for the acceptance of such sentences. thus, in typical situations where reciprocal sentences like (1) and (4) are considered true, the two corresponding binary sentences will be judged true as well, but there might be atypical situations where this judgement will not go through. most natural concepts have more than one typical attribute (e.g. a fruit is typically both round and sweet). we show that something similar holds of reciprocal verb concepts by identifying an attribute separate from symmetric participation that contributes to the typicality of reciprocal events. this attribute, which we call collective intentionality, describes the sum of relevant intentions and emotions of group members towards the event. in sam and julia hugged, we might expect sam and julia to be happy and affectionate. in sam and julia fought, sam and julia are likely to be angry and aggressive. both symmetric participation and collective intentionality are assumed to contribute to the typicality of reciprocal events. accordingly, we expect asymmetric participation to be more easily tolerated in case where the agents show the relevant kind of collective intentionality. for instance, we hypothesize that the possibility of judging (2) as false when (1) is true is boosted in situations where both sam and julia show affection to each other. in binary sentences where the agent is singular, we expect that only the agent’s intentionality affects the proceedings of elm 1: 197-203, 2021 imke kruitwagen, yoad winter and james hampton: reciprocal predicates: a prototype model. 198 https://doi.org/10.3765/elm https://www.elm-conference.net/ acceptability of the sentence, whereas the intentionality of the other participant (patient or theme) is assumed to have a smaller effect. the paper is structured as follows. section 2 reports an experimental study of two hypotheses: (i) symmetric participation is not a necessary condition for reciprocal verbs; and (ii) collective agent intentionality can boost the acceptability of sentences with reciprocal verbs, whereas patient/theme intentionality does not have an equal effect on the acceptance of binary sentences. section 3 discusses the results and proposes an account in terms of hampton’s (2007) threshold model. 2. experiment. in the reported experiment, we collected truth-value judgements on sentences with unary and binary forms. dutch speakers were shown situations where symmetric participation is likely to be missing, and which varied with respect to collective intentionality. each experimental item consisted of a short video clip and a dutch sentence, either a unary sentence of the form a and b verb (e.g. “violet and mark fought”) or a binary sentence of the form b verb (preposition) a (e.g. “mark fought (with) violet”). all videos depict situations with two characters, one of whom visibly performs the relevant physical action on the other character, while the other character remains passive. for instance, for the verb knuffelen (“hug”), we used video clips that showed a woman (“violet”) clasping a man in her arms, while the man (“mark”) does not clasp her in his arms. binary sentences (e.g. “mark hugged violet”) all contained a subject that refers to the passive character, and are therefore expected to be judged as false. if sentences with unary forms are accepted despite rejection of the corresponding binary sentences, we conclude that the unary form does not require symmetric participation. the video clips that were used were of two types. in one type, both characters showed the intention that is intuitively expected with respect to the physical act (e.g. affection for hug, anger for fight). we consider these video clips as illustrating collective intentionality. in the other type of video, only the active character showed the expected intention, whereas the reaction of the passive character was unsuitable to the act (e.g. disgust for hug, indifference for fight). both types of video were designed as lacking the relevant physical act of the “passive” character. thus, if sentences with the unary form are accepted significantly more in the situations with collective intentionality despite equal acceptance of the binary sentence, we conclude that collective intentionality boosts the acceptance of sentences with unary reciprocal forms. 2.1. method. 2.1.1. participants. a total of 449 participants (287 female, age m = 19) took part in the experiment. they were all enrolled in a first year bachelor’s course at utrecht university and they took part in the experiment as part of a class. participants did not receive monetary compensation. 2.1.2. materials. the unary and binary forms of four dutch reciprocal verbs were tested for their acceptability, each of them in two different situations: with/without collective intentionality. in total there were 16 test items, each of which consisting of a short video clip and a dutch sentence. in addition, participants were asked about their age, gender and native language. target item sentences were constructed from verbs that have both a binary form (e.g. “a hugged b”) and a unary form with a collective interpretation (e.g. “a and b hugged”). in order to optimally represent the semantic variety of reciprocal verbs, these four verbs were selected: • knuffelen (hug) • botsen (tegen) (collide (with)) • vechten (tegen) (fight (against)) proceedings of elm 1: 197-203, 2021 imke kruitwagen, yoad winter and james hampton: reciprocal predicates: a prototype model. 199 https://doi.org/10.3765/elm https://www.elm-conference.net/ • fluisteren (tegen) (whisper (to)) unary and binary target sentences were included for each verb. in addition to the 8 target item sentences, 8 filler item sentences were constructed. the filler items served to obtain a better balance in the expected number of true/false reactions and to prevent the participants from using automatic strategies when reacting to the target stimuli. the filler item sentences included both unary and binary verbs. for each target verb, two different videos were used. the videos were approximately 30 seconds long, showing a woman (“violet”) and a man (“mark”). in all target videos, the woman carries out the action described by the verb. for instance, in case of hug, she wraps her arms around the man; in case of talk, she speaks. the man remains passive and does not physically carry out the physical actions described by the verb. in case of hug, for instance, the man does not wrap his arms around the woman; in case of talk, he does not speak. the difference between the two videos for each verb is in the intentionality of the man. one type of videos, henceforth “ci videos”, demonstrated collective intentionality in that the man uses social cues such as facial expressions and body language to show his involvement in the action. the second type of videos, “no-ci videos”, is similar to the ci videos except for one respect: the man now expresses a negative attitude towards the action or is indifferent towards it. we used a between-subject design with a total of eight versions. each version contained three items: a target item with a unary sentence, a target item with a binary sentence, and a filler item. the verbs in the three items were all different. four versions contained only target items with ci videos, while the four other versions only contained target items with no-ci videos. this was done to prevent participants from consciously noticing a difference in the man’s involvement towards the action, which could influence their judgements. 2.1.3. procedure. the experiment was designed using limesurvey and an application developed at utrecht university with the aim of randomizing the allocation of participants to questionnaires (https://rocky.sites.uu.nl/tools/). the experiment was run in a lecture room during the break of a lecture. the url to the experiment was presented on a large screen. following the url led participants to a randomized version of the experiment. participants used earphones and individual laptop, tablet or smart phone. for each item, a video was displayed, after which the accompanying sentence was shown on the screen. participants were asked to indicate whether they judged the sentence as “true” or “false” in the preceding video. they were told that there are no right or wrong answers, but that we as researchers were curious about their intuition. they were also instructed to not think too long about their judgement. 2.1.4. analysis. fifteen participants were excluded from analysis because they were not native speakers of dutch. for each of the target items, the proportion of “true”-responses was computed. then, a logistic regression model was constructed, using sentence type (unary or binary) and video intention (ci or no-ci) as predictor variables. 2.2. results. the results demonstrate a difference in acceptance rates between all four target types, as shown on figure 1. proceedings of elm 1: 197-203, 2021 imke kruitwagen, yoad winter and james hampton: reciprocal predicates: a prototype model. 200 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: results experiment with respect to the ci videos, there is a substantial difference between acceptance rates on the unary and binary forms. for unary forms with a ci video the acceptance rates varied from 40% for fight to 90% for whisper, while the acceptance rates for binary forms on ci videos varied from 4% for collide to 39% for hug. furthermore, there is also a substantial difference between acceptance rates on unary forms in ci videos and unary forms in no-ci videos. acceptance rates for unary sentences in no-ci videos varied from 18% for fight to 69% for collide. however, for collide there was no difference between the acceptance rates on the unary sentence with a ci video and the unary sentence with a no-ci video. logistic regression analyses were computed, using sentence type (unary vs. binary), intention (ci vs. no-ci) and sentence type × intention interaction as predictor variables. separate analyses were done for each of the four verbs. for all verbs a significant effect of sentence type was found. thus, sentences with the unary form of the verb yielded significantly more true answers when compared to sentences with the binary form. for all verbs except collide there was a significant effect of intention on the acceptance of both unary and binary form sentences. thus, except for collide, unary and binary form sentences combined with the ci videos were accepted more frequently than unary and binary form sentences combined with no-ci videos. no verb exhibited a significant interaction effect between intention and sentence type. thus, there was no statistical evidence that intention had a larger effect on acceptance of unary forms than on binary forms. 3. discussion. previous proposals on the semantics of reciprocal alternations assume symmetric participation: a logical entailment from sentences with the unary form (“violet and mark hugged”) to their binary counterparts (“violet hugged mark” and “mark hugged violet”). while this entailment is very likely to hold for reciprocal verbs like marry and date, our work questions it for verbs like hug or collide. as an alternative, the following hypotheses were tested: (i) symmetric participation is not a necessary condition for reciprocal verbs; and (ii) collective intentionality can boost the acceptability of sentences with unary form reciprocals, beyond its possible effect on the acceptance of binary form sentences. our results showed situations where unary sentences yield 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% whisper hug fight collide acceptance rates experiment unary + ci binary + ci unary ci binary -ci proceedings of elm 1: 197-203, 2021 imke kruitwagen, yoad winter and james hampton: reciprocal predicates: a prototype model. 201 https://doi.org/10.3765/elm https://www.elm-conference.net/ significantly more true answers compared to their binary correlates. thus, while many participants saw that the event lacked symmetric participation, this did not prevent many participants from judging the unary sentences as true. it is unlikely that symmetric participation holds for all speakers. this supports hypothesis (i) and is evidence against the entailment theory. as for our second hypothesis, with 3 of the 4 verbs that were tested (hug, fight and whisper), collective intentionality significantly boosted the acceptance of both unary and binary sentences. there was no statistical evidence that collective intentionality had a larger effect on the acceptance of unary form reciprocals than on binary forms. hypothesis (ii) is thus only partially supported. for collide there was no effect of collective intentionality at all. this is not surprising: typically, there is no intentionality involved in an event of colliding. for the other verbs, it is tempting to conclude that collective intentionality is a typical feature for both the unary and the binary from of reciprocal verbs. thus, surprisingly, the intentionality of a non-agent participant might matter as much as that of the agent’s. if this is the case, effects of collective intentionality on acceptance of reciprocal verbs may only be an epiphenomenon of its effect on binary forms: the acceptability of unary form sentences may depend only on the acceptability of their binary correlates. under this interpretation, since collective intentionality positively influences acceptance of the binary form, it indirectly affects acceptance of the unary form as well. this line of explanation is in line with gleitman et al.’s (1996) proposal that binary forms of (at least some) reciprocal verbs have typical symmetry about them, even if the symmetric participation that gleitman et al. assume for reciprocal entries is not supported by our results. our experimental results lead to a new proposal regarding the reciprocal alternation. instead of the logical entailment that is involved in the idea of symmetric participation, we propose a typicality approach, where conceptual features connect the unary and binary form. the conceptual features that are shared by all reciprocal verbs are symmetric participation and collective intentionality. the combination of the verb form (unary or binary) and the conceptual content of its stem determines the typicality of different situations for sentences made of the verb. if one conceptual feature has a non-typical value (in our case, lack of symmetric participation), the other conceptual feature (in our case, collective intentionality) can boost the overall typicality of the situation. in such cases the event may end up passing the threshold of belonging to the relevant category of events despite the lack of one of the typical features as in hampton’s (2007) threshold model. consequently, the corresponding sentence may become acceptable. different verb concepts may have different weights assigned to relevant features. for example, for collide, collective intentionality is not an important conceptual feature: the contribution of collective intentionality to event typicality is low, or even zero. by contrast, for whisper, collective intentionality is proposed to be an important feature, which explains why the acceptance of “violet and mark whispered” increases significantly when violet and mark show collective intentionality, compared to a minimally different event in which mark does not show intention. a remaining question that our study leaves open is to what extent collective intentionality is specifically a typical property of unary forms, and to what extent it also characterizes binary forms that participate in reciprocal alternations. references carlson, greg. 1998. thematic roles and the individuation of events. events and grammar, 35– 51. springer. https://doi.org/10.1007/978-94-011-3969-4_3 proceedings of elm 1: 197-203, 2021 imke kruitwagen, yoad winter and james hampton: reciprocal predicates: a prototype model. 202 https://doi.org/10.3765/elm https://www.elm-conference.net/ dimitriadis, alexis. 2008. irreducible symmetry in reciprocal constructions. in ekkehard könig & volker gast (eds), reciprocals and reflexives: theoretical and typological explorations. trends in linguistics 192, 375-409. berlin/new york: mouton de gruyter. https://doi.org/10.1515/9783110199147.375. dowty, david. 1991. thematic proto-roles and argument selection. language, 547-619. https://doi.org/10.2307/415037. gleitman, lila r. 1965. coordinating conjunctions in english. language, 41(2), 260-293. https://doi.org/10.2307/411878. gleitman, lila. r., harry gleitman, carol miller & ruth ostrin. 1996. similar, and similar concepts. cognition, 58(3), 321-376. https://doi.org/10.1016/0010-0277(95)00686-9. hampton, james. 2007. typicality, graded membership, and vagueness. cognitive science, 31(3), 355-384. https://doi.org/10.1080/15326900701326402. lakoff, george, and stanley peters. 1966. phrasal conjunction and symmetric predicates. report nsf-17, harvard computation lab. reprinted in reibel and schane ed., modern studies in english, prentice hall. 1969. siloni, tal. 2008. the syntax of reciprocal verbs: an overview. in ekkehard könig & volker gast (eds), reciprocals and reflexives: theoretical and typological explorations. trends in linguistics 192, 451-498. berlin/new york: mouton de gruyter. 451498. https://doi.org/10.1515/9783110199147.451. proceedings of elm 1: 197-203, 2021 imke kruitwagen, yoad winter and james hampton: reciprocal predicates: a prototype model. 203 https://doi.org/10.3765/elm https://www.elm-conference.net/ new data on the competition between definites and indefinites nadine bade & florian schwarz* abstract. in this paper, we report on four experiments investigating obligatory presupposition effects. specifically, we look at the inferences arising from not using presupposition triggers when their use is supported by the context. we compare these inferences and the contextual factors for their derivation to presuppositions and implicatures. extending previous work, we explore not only the english definite determiner “the” but also the dual “both” and their respective competition with the universal quantifiers “every” and “all”. keywords. presuppositions, non-presuppositions, implicatures 1. introduction. in this paper, we present a set of novel experimental data from english shedding light on the non-uniqueness effect associated with the indefinite determiner, as illustrated by the classic example in (1). (1) #a father of the victim arrived at the crime scene. (heim 1991) there is more than one father of the victim the oddness of (1) is usually attributed to the fact that indefinite determiner phrases (dp) cannot refer to unique objects due to a blocking effect evoked by the definite dp “the father of the victim”. the observation that definite marking must be used when it can be – given the presupposition of uniqueness is fulfilled – has been accounted for by postulating a general principle maximize presupposition, which has received substantial attention in the recent literature (heim 1991, chemla 2008, percus 2006, sauerland 2008, singh 2011, schlenker 2012, marty 2017, anvari 2018, spector & sudo 2017, rouillard & schwarz 2017, marty & romoli 2020). in this paper, we look at inferences arising as a result of reasoning with maximize presupposition and compare them to presuppositions and implicatures in different experimental settings. to test the generality of the phenomenon exemplified by (1) we also look at the competition between quantifiers “all” and “every” with the presuppositionally stronger “both” and “the”, respectively, see (2-a) and (2-b). (2) a. john broke {#all / both} of his arms. b. {#every / the} sun is shining. 2. background. 2.1. theoretical background. the principle maximize presupposition (heim 1991) has been posited as a general pragmatic principle to account for the obligatory insertion of presupposition triggers: maximize presupposition (heim 1991) make your contribution presuppose as much as possible! *nadine bade, university of potsdam (nadine.bade@uni-potsdam.de) & florian schwarz, university of pennsylvania (florians@ling.upenn.edu). proceedings of elm 1: 015-026, 2021 c©2021 nadine bade and florian schwarz published by the lsa with permission of the author(s) under a cc by license. 15 https://doi.org/10.3765/elm https://www.elm-conference.net/ it explains why (3-b) is preferred over (3-a), given world knowledge that everyone has a unique (biological) father. (parallel explanations extend to the other examples.) (3) a. #a father of the victim arrived at the crime scene. b. the father of the victim arrived at the crime scene. heim (1991) argued that the maxim of quantity could not account for these cases under the theoretical assumption that (3-a) and (3-b) are contextually equivalent (=equally informative assuming the truth of the presupposition) but differ in the definedness conditions they introduce. postulating a separate principle for these phenomena also aligns with data suggesting that inferences evoked by presuppositional competition show behavior which is different from presuppositions and implicatures (sauerland 2008). specifically, these inferences (called presuppositional implicatures in the following) have been argued to have a weaker status (resist strengthening), and project at the same time. that is, the oddness effect of the presuppositionally weaker alternative is preserved under holes for presuppositions, such as negation, see (4). (4) a. #not all arms of john are broken. b. #i did not see a father of the victim. however, this line of argument has been challenged in the more recent literature. most of the challenging data discussed revolve around claims that strengthening of presuppositional implicatures is possible, and that their derivation is mandatory under certain circumstances (chemla 2008, magri 2009, marty 2017, elliott & sauerland 2019). (3-a) is a prominent example of this effect. the oddness of the sentence seems to result from the non-uniqueness inference being mandatory, and blind to common knowledge (singh 2011, magri 2009). based on these challenging data, presuppositional implicatures have been claimed to be derived by the same mechanism as implicatures. nonetheless, the theoretical status of these inferences still remains unclear. there is no agreement on whether implicatures and presuppositional implicatures should be treated completely on a par (see e.g. rouillard & schwarz (2017) for discussion). it has been claimed that, just like for implicatures, the presence and relevance of the alternative, the knowledge state of hearer and speaker, as well as ease of accommodation are crucial factors impacting whether presuppositional implicatures arise and are strengthened (chemla 2008, marty 2017, elliott & sauerland 2019). the goal of the experiments reported below was to test the predictions of different theories with regard to what role these factors play. 2.2. previous experimental work. several claims in the theoretical literature about the role of alternatives for implicatures have been supported by experimental results. both the presence and relevance of the alternative play a role in deriving implicatures (bott & chemla 2016, rees & bott 2018, degen & tanenhaus 2015). they have been argued to be one deciding factor for whether the processing of implicatures is delayed (huang & snedeker 2009, 2011), or immediate (grodner et al. 2010). furthermore, implicatures have not only been shown to differ from literal meaning in processing but also to differ from presuppositions (bill et al. 2018). there are fewer experimental investigations of the processing of presuppositional implicatures. most of the existing literature focuses on the difference between indefinite and definite determiners. kirsten et al. (2014) find in an eeg experiment that unmet non-uniqueness inferproceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 16 https://doi.org/10.3765/elm https://www.elm-conference.net/ ences evoke higher processing costs for the indefinite than unmet uniqueness presuppositions for the definite determiner. they attribute this difference to additional cognitive load demanded by the introduction of a new discourse referent with the indefinite. another set of reading time studies tested the use of definites versus indefinites referring to stereo-typically unique items in contexts in which they are typically unique (e.g. stove in a kitchen) or not (e.g. stove in an appliance store) clifton jr (2013). clifton found that interactions between contexts and determiner in reading times only emerged if the experiment involved a secondary arithmetic task, which is argued to lead to deeper processing resulting in participants forming a more complete situation model. mouse-tracking data from schneider et al. (2019) support a view where unmet uniqueness for the definite and unmet non-uniqueness for the indefinite do not evoke parallel processing behavior. these data also stress the relevance of the presuppositional alternative for the indefinite. further evidence for differences between definite and indefinite comes from an eye-tracking study reported in bade & schwarz (2019a). they suggest that, if the inference is drawn, nonuniqueness evokes different gaze patterns than those arising when deriving a uniqueness presupposition. specifically, in line with schneider et al. (2019)’s results, the indefinite seems to involve more consideration of the alternative. in contrast, an eye-tracking experiment reported in bade & schwarz (2019b) reveals that if the contextual conditions for drawing the non-uniqueness are fully met the processing differences between indefinites and definites disappear. the results suggest that these contextual conditions involve complete awareness of the definite alternative and its relevance. in the experiment, this was achieved by including a production task providing definite and indefinite as alternatives. a similar pattern is observed in a mouse-tracking study by (schneider et al. 2020). in the study, participants were also presented with a production task first. the results show that both determiners initiate immediate movements towards the target, with barely any differences between determiners for the relevant measures considered. again, this suggest that awareness of the alternative matters, in that it makes the patterns for the determiners indistinguishable. 3. experiments. we ran four experiments in total. a first goal was to establish contexts where inferences involving reasoning over presuppositional alternatives reliably arise. the second goal was to identify the factors that make these alternatives salient, and compare them to the ones of other scale types. 3.1. study one: picture selection. 3.1.1. aims. the aim of this study was to test whether presenting material as a dialog affects target choices of the two determiners differently. a weakness of previous studies was that they left the overall discourse situation implicit or underspecified. in the current experiment, we made it clear that speaker and hearer had knowledge of the situation, including the number of items talked about, and that thas number was relevant. 3.1.2. design and material. we used a simple 2x2 design with determiner and picture competition fully crossed. sentences with an indefinite or definite determiner such as given in (5) were used in a comic strip, see figure 1a and 2a. (5) {a/the} shirt in my closet has a stain on it. proceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 17 https://doi.org/10.3765/elm https://www.elm-conference.net/ there were two different pairs of pictures shown as competing options along this comic, see figures 1 and 2. the first competition was between a picture with a single stained shirt and a picture with three stained shirts. the second competition was between a single stained shirt picture and a picture with three shirts, one of which was stained. there were, moreover, filler items using similar picture types but containing sentences with “only” or plurals with negation in the comic strip (the same fillers were used for all experiments). (a) sample comic with “the” (b) 3 shirts 3 stains (c) 1 shirt 1 stain figure 1: 3shirts-1stain versus 3shirts-1stain (a) sample comic with “a” (b) 3 shirts 1 stain (c) 1 shirt 1 stain figure 2: 3shirts-3stains versus 3shirts-1stain participants’ task was to pick the picture they thought matched the comic strip they saw. 3.1.3. predictions. sentences with an indefinite such as ”a shirt in my closet has a stain on it” should come with an inference of non-uniqueness (’there is not exactly one shirt in the closet’). accordingly, we predicted that pictures with non-unique objects would be targets for indefinites. conversely, sentences with the definite should come with a presupposition of uniqueness (’there is exactly one shirt in the closet’). as a result, pictures with unique objects are predicted to be targets for definites. however, a further complication, attested in previous results, arises from an implicature concerning the number of items of which the predicate is true (’there is exactly one shirt with a stain on it’), which seems to introduce a certain amount of infelicity relative to pictures where it is not met. we therefore expect an interaction of pictures and determiner: whereas the definite determiner should exhibit high rates of choices of unique object-pictures in both conditions, which align both with its presupposition and this implicature, the presuppositional implicature introduced by the indefinite determiner competes with this implicature in the picture pairing in figure 1; but in the condition illustrated in figure 2, where both pictures satisfy the ’exactly-one’-implicature, participants should choose the 3shirts-1stain picture type as it is the only choice that matches the presuppositional implicature. 3.1.4. results. we used a generalized linear mixed effects model analysis to compare the rate of 1shirt-1stain type picture choices. the model had random slopes for participants and random intercepts for items. we find a significant interaction between picture type and determiner (see stats in table 1). in line with predictions, the choice for the unique object picture is at ceiling for the definite across competitors. for the indefinite, a clear choice for the 3shirt picture in line proceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 18 https://doi.org/10.3765/elm https://www.elm-conference.net/ with the presuppositional implicature only transpired when the implicature was met (with just one stained shirt), otherwise, the 1shirt-1stain picture was the predominant choice, presumably to avoid incompatibility with the one stain implicature. estimate std.err z value p value 3shirts-3stains vs 3shirts 1stain:the -3.5605 0.4067 -8.754 < 2e-16 *** table 1: output of glmer for interaction between picture context and determiner to see whether determiners differed in the rate of target choices for the picture condition where both pictures satisfied the implicature, we looked at the interaction between context and determiner with target coded as the dependent variable. we see no difference in target choices between determiners when the implicature was satisfied by both pictures (β̂ = –1.83, se= 1.112, z = –1.645, p= 0.3532). significance was calculated based on least square means using the emmeans package in r. the statistical results are summarized in figure 3. figure 3: experiment 1 – percentage of choices for 1shirt-1stain type picture by determiner and competitor picture in sum, we find evidence of both the presupposition of the definite and the presuppositional implicature of the indefinite impacting picture choice rates. the rate of target choices for the indefinite and definite is the same when no interfering implicature is at play. once it is, the impact of the presuppositional implicature becomes minimal, as a desire to avoid an incompatibility with the one-stain implicature seems to dominate choice patterns. 3.2. study two: acceptability rating. 3.2.1. aims. the aim of the second study was to directly evaluate the acceptability of the sentences in the various picture contexts provided to speakers to see whether any oddness arises, and to which degree, for violating the different inference types under scrutiny. 3.2.2. design. the design is 2x2 with the factors determiner and picture type fully crossed. the latter was a group factor, i.e. people overall saw only one picture type. we tested individual picture types – 1shirt-1stain, 3shirts-1stain, 3shirts-3stains – with one of the two determiners. we asked people how natural they find the sentence in the given picture context on a scale from 1–7 (completely natural). proceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 19 https://doi.org/10.3765/elm https://www.elm-conference.net/ predictions the acceptability of the 1shirt-1stain picture should be higher than the 3shirts-1stain picture for the definite, due to the latter violating the uniqueness presupposition, or alternatively requiring domain restriction to make the presupposition true. reversely, for the indefinite, the 3shirts-1stain picture should be more acceptable than the 1shirt-1stain picture, in light of its presuppositional implicature of non-uniqueness. given theoretical assumptions the 1shirt-1stain picture should furthermore be more acceptable with the definite than indefinite. we thus predicted an interaction. 3.2.3. results. assuming an ordinal scale, we used a cumulative link model analysis and the clmm function in r to analyze the acceptability rating data. we find an interaction between picture type and determiner, see figure 4 and stats in table 2, but only for the comparison between 1shirt1stain and 3shirts-1stain. however, numerical differences are very small. figure 4: experiment 2 – average acceptability by determiner and picture type estimate std.err z value p value 3shirts 1stain:the -1.145046 0.432405 -2.648 0.00809 ** table 2: output of clmm for interaction between determiner and picture looking at contrasts, we see a difference between determiners only for the 3shirts-1stain picture, see table 3. contrast estimate std.err z.ratio p value 3shirts 1stain a vs the 1.3027 0.335 3.893 0.0014** table 3: output of emmeans for pairwise contrasts between determiner for 3items 1 stain picture type in sum, we see no substantial differences between determiners, suggesting that in the contexts given, their use is equally acceptable. this contrasts with the picture selection data reported above, which suggest awareness of the inferences these determiners come with.1 3.3. study three: covered box. 1the results are not likely to be the result of a ceiling effect, as we observe rather low ratings for our complex filler items. proceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 20 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.3.1. aim. the goal of the third experiment was to see whether a covered box paradigm would change response patterns. in addition to changing the task for experiment three, we added a contextual manipulation to test whether increasing the relevance of uniqueness being true would boost choices based on presuppositional implicatures of the indefinite. 3.3.2. design and material. again, we use a 2x2 design with context and determiner being fully crossed. context was a group factor. in the first context, the question posed by the parent in the picture was more neutral, see (6-a). in the second context, the question was more specific to the situation, see (6-b). (6) a. context “sad”: why are you sad? b. context “why”: why are you not getting dressed? c. target: {a/the} shirt in my closet has a stain. both contexts were paired with an overt picture showing a single shirt, which is stained, and a picture which is partially covered by boxes and revealed nothing about shirts, see figure 5. (a) example of a picture with a covered box (b) 1 shirt 1 stain figure 5: picture condition for both contexts participants were instructed to choose the covered box picture if they thought the overt picture did not match the context. 3.3.3. predictions. the prediction was that the indefinite is affected more by the context manipulation than the definite, i.e. that there are more overt target choices with the “sad” context than the “why” context for ”a shirt”-sentences, but an equal amount of target choices for sentences containing ”the”. the result should be an interaction between context and determiner. 3.3.4. results. we analyzed the rate of overt picture choices using a generalized linear mixed effect model analysis. we find neither an effect of context, nor an effect of determiner for either of the context types. overall, we see no choices of the covered box picture based on the non-uniqueness inference of the indefinite, see figure 6. the rate of choices for the overt picture depicting a unique stained item, here a shirt, were at ceiling for both determiners and across contexts. 3.4. study 4: extension of the design to all/both. 3.4.1. aims. the aims of experiment 4 were two-fold. first, the goal was to extend the data on competition of the definite with indefinite to the competition between the dual “both” and universal quantifier “all”, as well as the competition of the definite with the universal quantifier “every”. second, the goal was to compare the non-duality and non-uniqueness presuppositional implicatures of universal quantifiers to indirect implicatures evoked by “not all”/“not every” (=“some”). proceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 21 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 6: experiment 3 – rate of overt picture choices by context and determiner 3.4.2. design and material. sentences used in the comic strip for experiment 4 contained “both” and “all” with and without negation, see (7-a) and (7-c). furthermore, we tested the competition between “every” and “the” with sentences such as (7-b). participants saw both competitors for each pair, but only one of the pairs. negation was also a group factor, i.e. participants saw “both” and “all” either always with or always without negation. (7) a. the shirts in my closet {all/both} have a stain on them. b. {every/ the} shirt in my closet has a stain on it. c. the shirts in my closet are not {all/both} stain-free. the positive sentences in (7-a) and (7-b) were only paired with one picture competition each. for the competition between dual and “all”, one picture made the presuppositional implicature of non-duality true, see figure 7b, the other made the presupposition of duality true, see figure 7a. (a) 2 shirts 2 stains (b) 3 shirts 3 stains figure 7: picture competition for “both” and “all” for the competition between definite and “every”, there was a picture in which 3 of 3 shirts were stained (making non-uniqueness true), and a picture where there was a single stained shirt (making uniqueness true), see figures 8a and 8b. for the experiment with negation, there are three individual pictures types, see 9. negated universal quantifiers have an indirect “some”-implicature as well as an non-duality presuppositional implicature (projected). “both” and “not both” sentences have a presupposition of duality. given these assumptions, the properties of pictures depicted in 9 with regard to the inferences associated with “not all” and “not both” are given in table 4 below. there were two relevant picture competitions, one that paired the 2shirts-1stain with the proceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 22 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) 1 shirt 1 stain (b) 3 shirts 3 stains figure 8: picture competition for “every” and “the” (a) 3 shirts 0 stains (b) 2 shirts 1 stain (c) 3 shirts 1 stain figure 9: picture types paired for “not all” and “not both” sentences 2shirts-1stain 3shirts 1stain 3shirts 0stains 1) not both (impl./pres.) true/true true/false false/false 2) not all (impl./pres. impl.) true/false true/true false/true table 4: status of inferences associated with “not all” and “not both” for each picture type 3shirts-1stain picture. this picture competition allowed us to compare violated presuppositional implicature with violated presupposition. the indirect implicature is true in both cases. the second pairing is a 2shirts-1stain versus 3shirts-0stain picture pairing. it compares violated presuppositional implicatures of non-duality (former) with violated indirect implicature (latter). the two pairings appeared with both type of determiners. 3.4.3. predictions. we predicted choices to be driven by the presuppositional implicatures of non-duality and non-uniqueness associated with universal quantifiers. for both, participants should choose the target with three of three stained shirts. the definite’s presupposition should drive people’s choice to the picture with a unique stained shirt. the presupposition of “both” is predicted to make people choose the picture with exactly two shirts, both stained. regarding negation, we predicted a similar pattern. the projecting presuppositions and presuppositional implicatures should make the 3shirts-1stain picture the target for “not all”, and the 2shirts-1stain picture the target for “not both”, respectively. if these scales work similarly to the one containing {a, the}, we predict there to be more choices of the non-duality violating picture for “not both” (2shirts-1stain) given that the competitor (3shirts-0stain) violates the indirect “some”-implicature. 3.4.4. results. the findings for the cases without negation are exactly as predicted. for both “all” and “every” we see choices based on their presuppositional implicatures, both significantly differ for the choices for the definite and “both”, see table 5. the results look much more reliable than for indefinites, with target choices being at ceiling, proceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 23 https://doi.org/10.3765/elm https://www.elm-conference.net/ estimate std.err z value p value det both -5.6115 0.8085 -6.941 3.90e-12 *** det the 21.207454 0.001823 11636 <2e-16 *** table 5: output of glmer for main effect of determiner, in the case of “both” compared to “all”, for “the” compared to “every” see figures 10a and 10b. for the cases involving negation, we see the predicted interaction, see table 6. estimate std.err z value p value 2shirts 1stain vs. 3shirts-0stains:notboth 7.9661 1.0573 7.534 4.91e-14 *** table 6: output for glmer for interaction between determiner and picture pairing this interaction is driven by the fact that when the implicature is true, choices for the universal quantifier are driven by non-duality, and choices for “both” driven by duality. however, in the pairing where one picture is implicature violating, participants choose the 2shirts-1stain picture with both “not all” and “not both”, irrespective of the non-duality presuppositional implicatures being violated, see figure 10c. (a) percentage of choices for picture with exactly one shirt by determiner (b) percentage of choices for picture with exactly two shirts by determiner (c) percentage of 2 shirts 1 stain picture by determiner and competitor picture figure 10: results experiment 4 4. discussion. to sum up, non-uniqueness inferences associated with indefinite determiners drive picture choices (exp 1), but only when its ’exactly-one’ implicature is true (exp 1). this contrasts with previous findings where choices were not relying on non-uniqueness inferences to the same degree (bade & schwarz 2019a). the difference between the studies lay predominantly in the discourse situation being much more specified in the current experiments. we take this to suggest that the knowledge state of speaker and hearer as well as relevance of the alternative play a crucial role in deriving presuppositional implicatures. our results thus further stress the importance of making the alternative with the definite competitor salient. however, we see no effect of non-uniqueness in acceptability (exp 2) or when the competitor picture is a covered box, even with additional contextual pressure (exp 3). the findings suggest that both determiners are felicitous choices both in presupposition and presuppositional implicature violating contexts. these findings stand in contrast with previous observations that both violations give rise to pragmatic oddness (e.g. in the cases of “a father of the victim...” or “a proceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 24 https://doi.org/10.3765/elm https://www.elm-conference.net/ highest mountain...”). they suggest that the definite alternative is not automatically activated (e.g. by the lexicon), even in a context that makes it salient, but must be raised by an overt picture. our findings also raise a methodological issue: in how far did participants perceive the inferences to be contextually entailed with pictures? for the competition between “every” vs. “the” and “both” vs. “every” (exp 4) we see a clear effect of non-uniqueness and non-duality inferences with both universal quantifiers in picture choices. this contrasts with the less reliable findings for the indefinite. it suggests that the ambiguity of the indefinite may play a role, and different usages may associated with different sets of competitors, depending on what is at issue. a parallel we observe is that, as for indefinites, the choice falls on the presuppositional implicature violating picture (only) if the competitor is implicature violating. why one would take precedence over the other is an open question for any theory, but especially under those theories that propose unique operator analyses where implicatures and presuppositional implicatures are derived by the same mechanism (marty 2017). if this approach is adopted more needs to be said about the relevance of the presuppositional alternative, when it must be activated or can be ignored. further research is needed complementing the methodology used here to understand this relation between focus, competition and alternatives in the domain of presupposition versus assertion better. references anvari, amir. 2018. logical integrity. in proceedings of semantics and linguistic theory 28, vol. 28, 711. linguistic society of america. 10.3765/salt.v28i0.4419. bade, nadine & florian schwarz. 2019a. an experimental investigation of antipresuppositions. in ava creemers & caitlin richter (eds.), proceedings of penn linguistics colloquium 42, 31–40. bade, nadine & florian schwarz. 2019b. (in-)definites, (anti-)uniqueness, and uniqueness expectations. in proceedings of cogsci 2019, 119–125. bill, cory, jacopo romoli & florian schwarz. 2018. processing presuppositions and implicatures: similarities and differences. frontiers in communication 3. 44. bott, lewis & emmanuel chemla. 2016. shared and distinct mechanisms in deriving linguistic enrichment. journal of memory and language 91. 117–140. chemla, emmanuel. 2008. an epistemic step for antipresuppositions. journal of semantics 25(2). 141–173. clifton jr, charles. 2013. situational context affects definiteness preferences: accommodation of presuppositions. journal of experimental psychology: learning, memory, and cognition 39(2). 487. degen, judith & michael k tanenhaus. 2015. processing scalar implicature: a constraint-based approach. cognitive science 39(4). 667–710. elliott, patrick & uli sauerland. 2019. ineffability and unexhaustification. in m.teresa espinal, elena castroviejo, manuel leonetti, louise mcnally & cristina real-puigdollers (eds.), proceedings of sinn und bedeutung 23, 399–412. grodner, daniel j., natalie m. klein, kathleen m. carbary & michael k. tanenhaus. 2010. ”some,” and possibly all, scalar inferences are not delayed: evidence for immediate pragproceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 25 https://doi.org/10.3765/elm https://www.elm-conference.net/ matic enrichment. cognition 116(1). 42 – 55. heim, irene. 1991. artikel und definitheit. in arnim von stechow & dieter wunderlich (eds.), semantics: an international handbook of contemporary research, 487–535. berlin: mouton de gruyter. huang, yi ting & jesse snedeker. 2009. online interpretation of scalar quantifiers: insight into the semantics–pragmatics interface. cognitive psychology 58(3). 376 – 415. huang, yi ting & jesse snedeker. 2011. logic and conversation revisited: evidence for a division between semantic and pragmatic content in real-time language comprehension. language and cognitive processes 26(8). 1161–1172. kirsten, mareike, sonja tiemann, verena c. seibold, ingo hertrich, sigrid beck & bettina rolke. 2014. when the polar bear encounters many polar bears: event-related potential context effects evoked by uniqueness failure. language, cognition and neuroscience 29(9). 1147–1162. magri, giorgio. 2009. a theory of individual-level predicates based on blind mandatory scalar implicatures. natural language semantics 17(3). 245–297. marty, paul. 2017. implicatures in the dp domain: massachusetts institute of technology dissertation. marty, paul & jacopo romoli. 2020. presuppositions, implicatures, and contextual equivalence. natural language semantics https://semanticsarchive.net/archive/ tg2nzkym/contextual.pdf. percus, orin. 2006. antipresuppositions. in aymui ueyama (ed.), theoretical and empirical studies of reference and anaphora : toward the establishment of generative grammar as empirical science, japan society for the promotion of science. rees, alice & lewis bott. 2018. the role of alternative salience in the derivation of scalar implicatures. cognition 176. 1–14. rouillard, vincent & bernhard schwarz. 2017. epistemic narrowing for maximize presupposition. in andrew lamont & katerina a. tetzloff (eds.), proceedings of north east linguistic society 47, 49–62. sauerland, uli. 2008. implicated presuppositions. in sentence and context, mouton de gruyter. schlenker, philippe. 2012. maximize presupposition and gricean reasoning. natural language semantics 20 (4). 391–429. schneider, cosima, nadine bade, michael franke & markus janczyk. 2020. presuppositions of determiners are immediately used to disambiguate utterance meaning: a mouse-tracking study on the german language. psychological research 1–19. schneider, cosima, carolin schonard, michael franke, gerhard jäger & markus janczyk. 2019. pragmatic processing: an investigation of the (anti-)presuppositions of determiners using mouse-tracking. cognition 193. 104024. https://doi.org/10.1016/j.cognition.2019.104024. singh, raj. 2011. maximize presupposition! and local contexts. natural language semantics 19. 149–168. spector, benjamin & yasutada sudo. 2017. presupposed ignorance and exhaustification: how scalar implicatures and presuppositions interact. linguistics and philosophy 40(5). 473–517. 10.1007/s10988-017-9208-9. proceedings of elm 1: 015-026, 2021 nadine bade and florian schwarz: new data on the competition between definites and indefinites. 26 https://doi.org/10.3765/elm https://www.elm-conference.net/ the role of relevance, competence, and priors for scalar inferences polina tsvilodub, bob van tiel, & michael franke* abstract. although it is often assumed that the natural language expressions ‘some’ and ‘or’ are interpreted according to their first-order logic counterparts, in certain contexts, they receive a narrower interpretation: ‘some’ is strengthened to ‘some, but not all’, and ‘or’ to ‘or, but not both’. this process is typically explained as an instance of scalar inference. to test this scalar inference hypothesis, we collect experimental evidence for the effects and interactions of three factors that have been argued to affect the robustness of the scalar inferences of ‘some’ and ‘or’: the relevance of the stronger alternative, the speaker’s competence about the alternative, and the prior probability that the alternative is true. we find that the interpretation of both triggers was affected by speaker competence, but only the interpretation of ‘some’ was also affected by prior probability, while relevance did not affect the interpretation of either trigger. ultimately, our results suggest that the interdependence of the three factors is more complex than just the sum of their effects. keywords. scalar inference; relevance; competence; prior probability; disjunction 1. introduction. it is often assumed that the natural language expressions ‘some’ and ‘or’ are equivalent to the existential quantifier ∃ and inclusive disjunction ∨ from first-order logic (e.g., grice 1975). however, in certain contexts, ‘some’ and ‘or’ appear to receive a narrower interpretation than their alleged logical equivalents. for example, neither of the intuitive inferences in (1) and (2) follow from the proposed logical equivalence. (1) peter ate some of the doughnuts. ⇝ peter did not eat all of the doughnuts. (2) peter ate a doughnut or a beignet. ⇝ peter did not eat both a doughnut and a beignet. the narrowing inferences of ‘some’ and ‘or’ are usually assumed to be instances of a more general type of inference called scalar inference (e.g., horn 1972, gazdar 1979, geurts 2010). scalar inferences are so-called because they are associated with lexical scales consisting of expressions that are, inter alia, lexicalised to the same degree and ordered in terms of logical strength. in the case at hand, ‘some’ is said to be associated with the scale ⟨some, all⟩ and ‘or’ with ⟨or, and⟩. (positive) utterances containing a lower-ranked scalar expression may imply that the corresponding sentence with the higher-ranked expression is false. scalar inferences are often explained as a variety of conversational implicature. conversational implicatures are inferences that can be explained on the basis of an argument that revolves around the assumption that the speaker is cooperative. grice (1975) develops the notion of cooperativity by arguing that cooperative speakers tend to follow certain conversational maxims. for *many thanks to natalie clarius, neele witte and elisa kreiss for help in creating materials and to stela ilieva and malin spaniol for help with implementation during early stages of this project. authors: polina tsvilodub, osnabrück university (polina.tsvilodub@gmail.com), bob van tiel, radboud university nijmegen (bobvantiel@gmail.com) & michael franke, university of tübingen (mchfranke@gmail.com). proceedings of elm 2: 288-298, 2023 c©2023 polina tsvilodub, bob van tiel michael franke published by the lsa with permission of the author(s) under a cc by license. 288 https://doi.org/10.3765/elm https://www.elm-conference.net/ example, cooperative speakers tend to be truthful (quality) and informative (quantity). based on these two maxims, the scalar inference in (1) can be explained as follows: 1. the speaker said ‘peter ate some of the doughnuts’. 2. she could have said ‘peter ate all of the doughnuts’ (maxim of quantity). 3. this would have been more informative, and hence cooperative. 4. so why didn’t the speaker utter the more informative alternative? 5. presumably, the speaker does not believe that the alternative is true (maxim of quality). 6. it is likely that the speaker knows whether the alternative is true or false. 7. hence, the speaker believes that the alternative is false. the same argument, mutatis mutandis, can be given to explain the scalar inference of ‘or’, as exemplified in (2). in that case, the relevant alternative is ‘peter ate a doughnut and a beignet’, and the resulting inference says that the speaker believes peter did not eat both a doughnut and a beignet. we will call this implicature-based account of the scalar inferences of ‘some’ and ‘or’ the standard account. while the standard account is almost universally adopted for ‘some’, its application to ‘or’ has been more controversial (e.g., simons 2001, geurts 2006, zondervan 2010). in particular, the use of ‘or’ tends to trigger the inference that the speaker is unsure which of the two disjuncts is true (e.g., whether peter ate a doughnut or a beignet). this ignorance inference makes it a priori unlikely—though not impossible—that the speaker knows whether or not the disjuncts may be jointly true. zondervan (2010) calls this the speaker expertise paradox. in any case, according to the standard account, scalar inferences are inferences to the best interpretation (atlas & levinson 1981). the hearer interprets the speaker’s utterance by trying to explain why the speaker uttered one sentence rather than a more informative alternative. in the case at hand, the most plausible hypothesis is assumed to be that the speaker believes that the alternative is false. but it is part and parcel of inferences to the best interpretation that, depending on the context, different explanations for the speaker’s behaviour may emerge as the most plausible. for example, the robustness of scalar inferences has been argued to be influenced by competence, relevance, and prior probabilities. in the next section, we discuss these three factors in more detail. after that, we describe our experiment in which we tested the effects of these three factors on the robustness of the scalar inferences of ‘some’ and ‘or’. 2. factors influencing the robustness of scalar inferences. 2.1. competence. the derivation of scalar inferences is often presented as a two-step process. first, it is inferred that the speaker avoided producing the alternative because she does not believe that the alternative is true. second, this weak inference can be strengthened to the scalar inference that the speaker believes the alternative to be false. crucially, the step from the weak inference to the scalar inference relies on the assumption that the speaker knows whether or not the alternative is true. this assumption is called the competence assumption (step 6 above) (e.g., sauerland 2004, soames 1982). goodman & stuhlmüller (2013) experimentally investigated the effect of the plausibility of the competence assumption on the robustness of the scalar inference of ‘some’. in their exp. 1, participants were presented with vignettes in which a speaker uttered sentences containing the trigproceedings of elm 2: 288-298, 2023 polina tsvilodub, bob van tiel michael franke: the role of relevance, competence, and priors for scalar inferences. 289 https://doi.org/10.3765/elm https://www.elm-conference.net/ ger ‘some’. the vignettes were designed so as to vary with respect to the assumed knowledge of the speaker. in the full knowledge condition, the speaker was shown to be competent; in the partial knowledge condition, the competence assumption was not satisfied. goodman & stuhlmüller found that participants in the partial knowledge condition were significantly less likely to derive the scalar inference compared to the full knowledge condition. the effect of the plausibility of the competence assumption on the scalar inference of ‘or’ has not been studied experimentally before. however, it has been observed that there is an apparent tension between the competence assumption for ‘or’ and its ignorance inferences. to illustrate, consider again (2). an utterance of this sentence tends to give rise to the ignorance inferences that the speaker does not know whether peter ate a doughnut, and she does not know whether peter ate a beignet. given these ignorance inferences, it is difficult to imagine that the speaker is confident that peter did not eat both a doughnut and a beignet, i.e., that the competence assumption is satisfied (geurts 2006, zondervan 2010). in the case at hand, that would require, e.g., that the speaker watched peter having lunch from afar, seeing that peter ate one and only one thing which the speaker could make out to be either a doughnut or a beignet. 2.2. relevance. an alternative explanation for the speaker’s decision to produce an informationally weaker utterance is that the speaker assumed that the hearer would not be interested in the added information expressed by the alternative. to illustrate, compare the following dialogues (from van kuppevelt 1996): (3) a: how many of the boys were at the party? b: some of the boys were at the party. (4) a: were some of the boys at the party? b: some of the boys were at the party. a’s question in (3) makes it clear that she is interested in the precise number of boys who were at the party. by contrast, a’s question in (4) suggests that she is only interested in whether or not some of the boys were at the party. in other words, the information that not all of the boys were at the party is intuitively more relevant in the first dialogue than in the second. consequently, in the second dialogue, b might have chosen to produce the informationally weaker ‘some’, not because she lacks evidence for the alternative containing ‘all’, but rather because she thought the hearer would have little or no interest in the extra information conveyed by the corresponding sentence with ‘all’ (the speaker might even consider the added information to be distracting for the hearer). in line with this observation, the scalar inference is intuitively more robust in the first dialogue compared to the second. in a series of experiments, zondervan (2010) investigated the effects of relevance on the robustness of the scalar inferences of ‘most’, which we may assume to pattern similarly to ‘some’ and ‘or’. zondervan constructed vignettes that made the corresponding scalar inferences either relevant or irrelevant, where relevance was manipulated in various ways (e.g., by means of explicit questions, prosodic emphasis, or contextual cues). participants were then presented with a statement containing the weaker term, even though the vignette made it clear that the statement with the stronger term was true. they had to indicate whether the statement was true or false, given the background story. zondervan consistently found that scalar inference rates (i.e., ‘false’ responses) proceedings of elm 2: 288-298, 2023 polina tsvilodub, bob van tiel michael franke: the role of relevance, competence, and priors for scalar inferences. 290 https://doi.org/10.3765/elm https://www.elm-conference.net/ were higher when the scalar inference was relevant than when it was not. however, the difference was typically rather small. for instance, in exp. 1, the scalar inference of ‘or’ was drawn less than 20% more often when it was relevant (73% of the time) than when it was not relevant (55%). 2.3. prior probability. the nature of the effect of prior probability on the robustness of scalar inferences is more contentious than the effects of competence and relevance. to illustrate, consider the following sentence (from geurts 2010): (5) cleo threw all her marbles in the swimming pool. some of them sank to the bottom. a priori, it is highly likely that all of the marbles sank to the bottom. how does this observation influence the robustness of the scalar inference? geurts (2010) intuits that the strength of the scalar inference is unaffected by the fact that one would naturally expect all of the marbles to sink. geurts’ intuition contrast with the predictions made by the rational speech act (rsa) model, a recent formalisation of the pragmatic reasoning process that underlies the derivation of scalar inferences (e.g., frank & goodman 2012). according to the rsa model, the prior probability of the stronger alternative should negatively correlate with the strength of the scalar inference, so that, in the example above, the scalar inference should be weak or even altogether absent. degen et al. (2015) experimentally investigated the effect of prior probability on the robustness of the scalar inference of ‘some’. in exp. 1, they presented participants with event descriptions such as ‘john threw 15 marbles into a pool’. these event descriptions were followed by a question such as ‘how many of the marbles sank?’. participants had to indicate the probability of each possible event (e.g., one marble sinking, two marbles sinking, and so on). in exp. 2, a different group of participants read the same event descriptions, but this time the descriptions were followed by a well-informed character producing an utterance like ‘some of the marbles sank’. participants again had to indicate the probability of each possible event, but this time based on the utterance rather than their prior expectation. degen et al. (2015) observed a correlation between prior probability and the strength of the scalar inference so that, when the ‘all’ situation was judged likely in exp. 1, it was also judged likely in exp. 2. 2.4. predictions. to sum up, based on literature, we hypothesize that scalar implicatures are sensitive to these three contextual cues in the following way: 1. competence: scalar inferences are more robust if the speaker is competent, i.e., knows whether or not the stronger alternative is true. 2. relevance: scalar inferences are more robust if the information expressed by the stronger alternative is relevant to the hearer with respect to the purpose of the conversation. 3. prior probability: scalar inferences are more robust if the information expressed by the stronger alternative is a priori likely to be false. in this paper, we systematically investigate these effects and their interactions on the robustness of the scalar inferences associated with ‘some’ and ‘or’. our goals are twofold. first, we aim to obtain evidence as to which contextual factors influence the robustness of a scalar inference. recent research has focused on variability in scalar inference rates across different scalar words (e.g., why the inference from ‘some’ to ‘not all’ is much more robust than the inference from ‘pretty’ to ‘not beautiful’, e.g., van tiel et al. 2016). here, we study variability in the robustness proceedings of elm 2: 288-298, 2023 polina tsvilodub, bob van tiel michael franke: the role of relevance, competence, and priors for scalar inferences. 291 https://doi.org/10.3765/elm https://www.elm-conference.net/ of scalar inferences that make use of the same lexical scales. by doing so, we obtain a more direct insight into the mechanism that underlies pragmatic inferencing, which in turn may also provide important new knowledge about the factors that cause cross-scalar variability. second, we seek to compare the scalar inferences of ‘some’ and ‘or’ with respect to their sensitivity to the three contextual factors. if these two types of scalar inference are indeed caused by the same underlying mechanism—as the standard account argues—it is natural to expect that they are similarly sensitive to the pragmatic factors under investigation. hence, if we observe marked differences in how competence, relevance, and prior probability affect the robustness of the scalar inferences of ‘some’ and ‘or’, that would provide at least circumstantial evidence in favour of the idea that they are aetiologically distinct, too. to address these goals, we conduct an experiment comparing the effects of these factors for both triggers. we describe this experiment in the next section. 3. experiment. to operationalize these factors experimentally, we designed context stories (vignettes) which we intuitively judged to score either high or low with respect to each factor of interest. that is, we manipulated the contextual relevance of the stronger alternative to the listener, the speaker’s competence about the truth of the statement with ‘all’ or ‘and’, and the prior probability of the statement with ‘all’ or ‘and’ to be true. we studied how these manipulations influenced the the robustness of the scalar inferences of ‘some’ (‘some but not all’) and ‘or’ (‘or, but not both’). the context stories and critical trials were designed in an analogous fashion for both triggers and varied within-subjects, so the following descriptions apply for both triggers. 3.1. materials and procedure. this study was a 2 × 2 × 2 × 2 within-subjects rating task (relevance × competence × prior × trigger), conducted as a web-based experiment.1 on critical trials, participants were asked to rate four sentences, one per factor, and one containing an upper-bounded (‘some’) or exclusive (‘or’) paraphrase of the trigger. on each trial, participants read a context story, followed by a sentence presented in a blue box which was meant to elicit the likelihood rating for a given factor. the sentence had either of the following forms: (6) relevance rating: it is important to x to know whether y. competence rating: z knows whether y. ‘or’ prior rating: if a, then b. / if b , then a. ‘some’ prior rating: if some y, then all y. inference strength rating: from what z said we may conclude that w. where x was the listener, z was the speaker, y was the target event in the background story and w was the event under the upper-bounded or exclusive reading. a and b were the disjuncts of y for ‘or’ vignettes. for the inference strength elicitation, participants additionally read a critical utterance of the form: (7) z says to x: y. which contained the trigger (‘some’ or ‘or’), presented below the background story in a red box, before rating the inference strength sentence. 1the experiment can be viewed at https://magpie-xor-some.netlify.app/. the preregistration for the experiment can be found at https://osf.io/v7zjp. proceedings of elm 2: 288-298, 2023 polina tsvilodub, bob van tiel michael franke: the role of relevance, competence, and priors for scalar inferences. 292 https://doi.org/10.3765/elm https://www.elm-conference.net/ below each sentence, participants were asked to indicate how likely it is that the statement in the blue box is true given the story, followed by a slider rating bar labeled ‘certainly false’ (left) and ‘certainly true’ (right). the slider positions were converted to 0–100 ratings. we designed 32 different stories per trigger type (‘some’ vs. ‘or’), resulting in a total of 64 stories. each participant saw eight randomly sampled stories (one per prior × competence × relevance condition out of four possible stories) such that they saw four ‘or’ and four ‘some’ stories in randomized order. the assignment of conditions to the triggers was randomized betweenparticipants. the experiment proceeded as follows (see fig. 1). first, participants were welcomed to the experiment and read instructions, which contained an annotated example to explain the meaning of the slider. figure 1: experiment procedure. the main part of the experiment consisted of eight critical vignettes, randomly shuffled with eight attention check vignettes. for each critical vignette, participants completed a block of 10 or 11 trials, consisting of a trial with a comprehension question, trials with statements eliciting relevance, competence and prior ratings (the last factor was elicited with two statements for the trigger ‘or’, see (6)), followed by three more trials with comprehension questions, and the critical scalar inference strength elicitation trial. comprehension trials were visually identical to the critical trials, but the statement to be rated only referred to the content of the background stories. they were designed so as to be either clearly true, clearly false or uncertain given the background story. the four comprehension statements for one vignette were sampled at random from six possible statements (two true statements, two false statements, and two uncertain statements). the attention checks consisted of one trial which visually matched the critical trials. on these trials, the vignettes contained a statement which wrote out what participants were supposed to answer (e.g., ‘please move the slider maximally left’) in an area of text that participants had to read to complete the trial. participants who failed more than two of these attention checks were excluded from the analysis. after the study, participants could voluntarily fill out a sociodemographic questionnaire. proceedings of elm 2: 288-298, 2023 polina tsvilodub, bob van tiel michael franke: the role of relevance, competence, and priors for scalar inferences. 293 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.2. participants. we recruited 277 participants through the crowd-sourcing platform prolific. participants were restricted to those whose first language included english, who previously took part in at least five other prolific studies, and whose approval rate was at least 0.9, according to prescreening criteria of the platform. participants took 20 minutes on average to complete the study and were compensated £2.48 for their participation. following our preregistered exclusion criteria, we excluded 15 participants for not indicating their native language, 3 participants for completing the study in under 8 minutes, 26 participants for failing more than two attention checks, and 27 for failing more than 20% of the comprehension questions. due to a coding error in the comprehension questions, 206 participants were left after applying exclusion criteria, although the preregistered target sample size after exclusions was 200 subjects. in the following analyses, data from the 206 participants is analysed.2 4. results. prior to conducting statistical analyses, we preprocessed the data by standardizing (z-scoring) the responses within each factor (relevance, competence, prior) for each participant. based on the aforementioned predictions, we expect that participants rate the scalar inference as more likely if (i) the alternative is rated as more relevant, (ii) the speaker is rated as being more competent, and (iii) the alternative is rated as a priori less likely. the remainder of this section is structured as follows: section 4.1 provides descriptive and confirmatory analyses following preregistration and section 4.2 provides additional exploratory analyses. 4.1. predictor ratings and confirmatory analyses. descriptively, participants’ factor ratings by-story agreed well with the designed classification of the stories (fig. 2, red vs. blue color on x-axis). that is, the ratings for the relevance, competence and prior statements were not distributed uniformly across the stories, but aligned with our prior categorisations (fig. 2, x-axis), validating our experimental manipulation of the explanatory factors. we analysed the ratings using a bayesian linear mixed effects model, regressing the target scalar inference ratings against the fixed effects of predictor ratings (i.e., the relevance, competence, and prior ratings elicited by the same participant for that vignette), the effect of trigger, and their two-, three-, and four-way interactions.3 we included random intercepts and random slope effects for the main effects of trigger, relevance, competence and prior by-subject, as well as random intercepts by-vignette.4 the categorical effect of trigger was dummy coded, using ‘some’ as the reference level. for all regression coefficients we used a wide and uninformative prior given by a t-distribution with mean 0, standard deviation of 2 and 1 degree of freedom. the model was fitted using the r brms package (bürkner 2017). we focus on the slope coefficients for the effects of relevance, competence and prior, once for ‘some’ and once for ‘or’. we check whether the posterior estimate of each effect was in either one of three intervals: (1) negative effect (≤ -0.05), (2) no effect (between -0.05 and 0.05), or 2all analyses were also conducted on data from 200 subject, excluding the last six submissions. no noteworthy quantitative or qualitative differences of the results were observed. 3the data and the analyses can be found under https://github.com/magpie-ea/ magpie-xor-experiment 4model in r syntax style: inference-rating ∼ trigger * relevance * competence * prior + (1 + trigger + relevance + competence + prior || subjectid) + (1 | vignette). for computational tractability reasons, the correlation of random effects by-subject was set to 0. proceedings of elm 2: 288-298, 2023 polina tsvilodub, bob van tiel michael franke: the role of relevance, competence, and priors for scalar inferences. 294 https://doi.org/10.3765/elm https://www.elm-conference.net/ relevance competence prior s om e o r −2 −1 0 1 2 −2 −1 0 1 2 −2 −1 0 1 2 −2 −1 0 1 2 −2 −1 0 1 2 participants' factor ratings p ar tic ip an ts ' i nf er en ce li ke lih oo d ra tin gs prior categorization of vignettes low high figure 2: ratings for relevance, competence and prior statements (x-axis) plotted against ratings for the strength of scalar inferences (y-axis). the top row shows ratings for ‘some’ (enriched to ‘some, but not all’). the bottom row shows ratings for ‘or’ (enriched to ‘or, but not both’). ratings for stories initially categorised (by the experimenters) as low (red) w.r.t. a given factor are on average lower (x-axis) than for those categorised as high (blue). (3) positive effect (≥ 0.05). we set the threshold for considering an effect as positive or negative to ±0.05 because we consider 0.05 to be the region of practical equivalence (rope) for the effect sizes we expect (kruschke 2014). we interpret the data as providing evidence in favor of an effect (positive, negative, no effect) if the posterior probability of the effect being true is ≥ 0.95 (i.e., 95% of posterior samples are in the corresponding interval). the probabilities of the respective coefficients lying in a particular interval are reported below (i.e., for instance, if p = 0.95 is reported, it means that 95% of the posterior samples of the given coefficient are in the respective interval). in particular, we speak of evidence in favour of the scalar inference account if the prior effect is negative, and the relevance and competence effects are positive. fig. 3a shows examples of simulated posterior distributions over effect size samples which would confirm all our hypotheses (for better visual comparison to the observed results shown in fig. 3b). consistent with predictions of the standard account, for the trigger ‘some’, we found a clear negative effect of prior probability of the stronger alternative ‘all’ being true, as indicated by the probability of the negative effect of prior being p = 0.999 (fig. 3b, prior (some), orange color). similarly, we found a clear positive effect for speaker competence (p = 1, fig. 3b, competence (some), green color). however, we did not find a clear effect of relevance (fig. 3b, relevance (some), split colors). if anything, the data supported the result that relevance may only marginally influence the robustness of the enriched interpretation. in conproceedings of elm 2: 288-298, 2023 polina tsvilodub, bob van tiel michael franke: the role of relevance, competence, and priors for scalar inferences. 295 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 3: distributions of the posterior samples of each factor of interest for the trigger ‘some’ and ‘or’. the colors and the vertical dashed lines indicate the effect intervals (from left to right: negative, no, positive effect). points indicate posterior means, thick lines indicate 66% credible intervals, thin lines indicate 95% credible intervals. a: idealized hypothetical results implied by the standard account (for comparison to the obtained empirical results). the effect size was set to 0.25 for visualization purposes. b: main experimental analysis results. negative no effect positive by-subject random slope some relevance 0.006 0.519 0.474 0.07 competence 0 0 1 0.13 prior 0.999 0.001 0 0.16 or relevance 0.001 0.204 0.795 0.07 competence 0 0.007 0.993 0.13 prior 0.635 0.357 0.008 0.23 table 1: results of the separate by-trigger exploratory models. columns indicate probabilities of respective effects being in the indicated range, except for the column ‘by-subject random slope‘ which shows estimates. trast, for the trigger ‘or’, we only found a positive effect of competence (p = 0.993, fig. 3b, competence (or), green color). we found no credible effects of the prior of the stronger alternative being true; results for relevance patterned with results for ‘some’ (fig. 3b, prior (or), relevance (or), split colors). therefore, the main analysis did not provide strong evidence in favor of the scalar inference account for ‘or’. comparing the overall results to ‘some’, they provide evidence against the identity hypothesis positing that the two triggers are interpreted via the same underlying mechanism. 4.2. exploratory analyses. the descriptive results visually suggested a possible effect of prior for ‘or’ (see fig. 2, lower right), which, however, was not borne out in the main analysis. to investigate the results for ‘or’ in more detail, we explored an analysis wherein the predictor ratings were averaged by-vignette and then regressed against the implicature strength ratings. this analysis amounts to averaging over the different participants and thereby removing possible by-subject proceedings of elm 2: 288-298, 2023 polina tsvilodub, bob van tiel michael franke: the role of relevance, competence, and priors for scalar inferences. 296 https://doi.org/10.3765/elm https://www.elm-conference.net/ variability.5 under this model, additionally to the main results, there was a credible effect of prior probability for ‘or’ (p = 0.999 of a negative effect), suggesting that by-subject effects might have overridden the main effect in the main analysis. following up, models for each trigger separately were fit (see tab. 1). those models included full random effects; the one for ‘or’ patterned with the main analysis, showing no credible prior effects (p = 0.635 of a negative effect). yet, the largest by-participant random slope was the by-subject estimate for the effect of prior, compared to the other factors. compared to the ‘some’ model, this slope estimate was also larger (tab. 1, last column). furthermore, an exploratory correlation analyses revealed a potential co-linearity between factors in the case of ‘or’ (r2 = -0.106 for relevance and prior, r2 = 0.127 for competence and relevance). no significant correlations were found for ‘some’. taken together, it seems that participants interpreted the prior statements for ‘or’ stories quite variably, which might be explained by differences in the prior expectations set up in the stories and whether these were contextual or based on world knowledge. this, in turn, might have influenced the (perceived) relevance to the listener and led to the observed correlated effects. future research should address these aspects more systematically. 5. discussion. our experiment provides novel results on the effects of the factors relevance, competence and prior on the interpretation of ‘some’ and ‘or’. however, future work may extend upon our results in several ways. first, our experiment considered how relevant the stronger alternative was to the hearer’s interests. yet other work rather focuses on relevance in terms of discourse purpose (e.g., van kuppevelt 1996), which could be formalized, e.g., in the form of explicit questions in the background stories. second, this paper looked at the derivation of the exclusive reading of ‘or’ through the lens of a scalar inference based account. however, alternative accounts derive the exclusive reading in terms of a distinctness condition, suggesting that disjunctions are infelicitous whenever the two disjuncts overlap, or either does not address the qud (e.g., simons 2001). especially for the latter point, the disjuncts need to be interpreted exhaustively, which results in the exclusive reading. our results provide no conclusive evidence with respect to this alternative account, calling for follow-up experiments manipulating exhaustivity and distinctness. beyond that, it will be interesting to determine which other factors—e.g., typicality (van tiel 2014), prosodic and linguistic prominence (breheny et al. 2006), and politeness (bonnefon et al. 2009)—might have similar effects. to sum up, our study provides new data on the effects of three contextual factors on the exclusive interpretation of ‘or’, and showed that they had a different effect than on the upper-bounded interpretation of ‘some’, calling into question the assumption that the interpretations of the two expressions are subject to the same pragmatic mechanisms. references atlas, jay david & stephen c. levinson. 1981. it-clefts, informativeness, and logical form: radical pragmatics (revised standard version). in peter cole (ed.), radical pragmatics, 1–62. academic press. bonnefon, jean-françois, aidan feeney & gaëlle villejoubert. 2009. when some is ac5model in r syntax style: inference-rating ∼ trigger * relevance * competence * prior proceedings of elm 2: 288-298, 2023 polina tsvilodub, bob van tiel michael franke: the role of relevance, competence, and priors for scalar inferences. 297 https://doi.org/10.3765/elm https://www.elm-conference.net/ tually all: scalar inferences in face-threatening contexts. cognition 112(2). 249–258. https://doi.org/10.1016/j.cognition.2009.05.005. breheny, richard, napoleon katsos & john williams. 2006. are generalised scalar implicatures generated by default? an on-line investigation into the role of context in generating pragmatic inferences. cognition 100(3). 434–463. https://doi.org/10.1016/j.cognition.2005.07.003. bürkner, paul-christian. 2017. brms: an r package for bayesian multilevel models using stan. journal of statistical software 80(1). 1–28. https://doi.org/10.18637/jss.v080.i01. degen, judith, michael henry tessler & noah d. goodman. 2015. wonky worlds: listeners revise world knowledge when utterances are odd. in proceedings of the 37th annual conference of the cognitive science society, 548–553. frank, michael c. & noah d. goodman. 2012. predicting pragmatic reasoning in language games. science 336. 998. https://doi.org/10.1126/science.1218633. gazdar, gerald. 1979. pragmatics: implicature, presupposition, and logical form. academic press. geurts, bart. 2006. exclusive disjunction without implicature. geurts, bart. 2010. quantity implicatures. cambridge university press. https://doi.org/10.1017/cbo9780511975158. goodman, noah d. & andreas stuhlmüller. 2013. knowledge and implicature: modeling language understanding as social cognition. topics in cognitive science 5. 173–184. https://doi.org/10.1111/tops.12007. grice, h. paul. 1975. logic and conversation. in peter cole & jerry l. morgan (eds.), syntax and semantics, volume 3: speech acts, 41–58. new york, ny: academic press. https://doi.org/10.1163/9789004368811 003. horn, laurence r. 1972. on the semantic properties of logical operators in english: university of california, los angeles dissertation. kruschke, john. 2014. doing bayesian data analysis: a tutorial with r, jags, and stan. academic press. van kuppevelt, jan. 1996. inferring from topics: scalar implicatures as topic-dependent inferences. linguistics and philosophy 19. 393–443. http://www.jstor.org/stable/ 25001633. sauerland, uli. 2004. scalar implicatures in complex sentences. linguistics and philosophy 27. 367–391. https://doi.org/10.1023/b:ling.0000023378.71748.db. simons, mandy. 2001. disjunction and alternativeness. linguistics and philosophy 597–619. https://doi.org/10.1023/a:1017597811833. soames, scott. 1982. how presuppositions are inherited: a solution to the projection problem. linguistic inquiry 13. 483–545. http://www.jstor.org/stable/4178288. van tiel, bob. 2014. embedded scalars and typicality. journal of semantics 31(2). 147–177. https://doi.org/10.1093/jos/fft002. van tiel, bob, emiel van miltenburg, natalia zevakhina & bart geurts. 2016. scalar diversity. journal of semantics 33. 137–175. https://doi.org/10.1093/jos/ffu017. zondervan, arjen. 2010. scalar implicatures or focus: an experimental approach: utrecht university dissertation. proceedings of elm 2: 288-298, 2023 polina tsvilodub, bob van tiel michael franke: the role of relevance, competence, and priors for scalar inferences. 298 https://doi.org/10.3765/elm https://www.elm-conference.net/ what’s the smallest part of spinach? a new experimental approach to the count/mass distinction sea hee choi & tania ionin* abstract. this paper reports on a study that uses a novel methodology, the minimal parts identification task, in order to probe the relationship between morphosyntax and interpretation. english, korean and mandarin chinese differ from one another with regard to the count/mass distinction. building on prior research, this study examines whether speakers of these three languages also differ in how they interpret count vs. mass nouns. the findings, while uncovering some language-specific effects of morphosyntax, point to the importance of universality, and suggest that interpretation drives morphosyntax rather than the other way around. keywords. atomicity; count/mass distinction; english; korean; mandarin chinese 1. introduction. the object/substance distinction is cognitive, while the count/mass distinction is linguistic; in plural-marking languages like english, there are a number of diagnostics for whether a noun is count or mass, see table 1. according to chierchia (1998a, b, 2010, 2015), the relevant semantic distinction underlying the count/mass morphosyntax is atomicity: a noun is atomic iff there exists a minimal unit that has the property denoted by the noun. thus, the minimal unit of chair is a chair, but there is no minimal unit of mustard. in languages like english, which have a fully grammaticized count/mass distinction, the relationship between atomicity and morphosyntax is not direct: e.g., furniture is atomic yet mass, while chocolate(s) can be either mass or count (see table 1). there is also cross-linguistic variation with regard to which nouns are count vs. mass: e.g., spinach is mass in english but count in french; beans is count in english but mass in russian. such nouns have been labeled flexible nouns in the literature: note that nouns can be flexible both across languages (as in the case of beans and spinach) and within a language (as in the case of chocolate(s) or stone(s) in english). diagnostic count nouns mass nouns indefinite article a (count) a chair / chocolate / bean *a furniture/mustard/spinach plural marking (count) chairs / chocolates / beans *furnitures/mustards/spinaches ability to occur in bare (determiner-less) form (mass) *i bought chair / bean. i bought furniture / mustard / spinach / chocolate. many (count) vs. much (mass) many chairs / chocolates / beans much furniture / mustard / spinach / chocolate table 1: diagnostics for the count/mass distinction in english unlike in plural-marking languages, in generalized classifier (gc) languages, where plural marking is optional, the relationship between atomicity and morphosyntax is direct. for example, in korean, only atomic nouns can combine with the plural marker -tul (kim 2005). in mandarin, the atomicity distinction is directly encoded in the classifier system (cheng and sybesma 1998). * thanks to the members of the illinois experimental linguistics group and to the elm audience for their feedback. this project was funded by an nsf dissertation research improvement grant, bcs-1823762.authors: sea hee choi, university of illinois at urbana-champaign (schoi76@illinois.edu) & tania ionin, university of illinois at urbana-champaign (tionin@illinois.edu). proceedings of elm 1: 113-124, 2021 c©2021 sea hee choi and tania ionin published by the lsa with permission of the author(s) under a cc by license. 113 https://doi.org/10.3765/elm https://www.elm-conference.net/ a number of experimental studies, beginning with barner and snedeker (2005), have investigated how speakers of both plural-marking languages (english, french) and gc languages (japanese, mandarin, korean) interpret different types of nouns. the present study follows in this tradition, but uses a novel methodology in order to directly examine which nouns are interpreted as atomic or non-atomic across different languages. we ask the following research question: does the morphosyntax of the count/mass distinction in a given language affect speakers’ interpretation of nouns as atomic vs. non-atomic? or, is interpretation driven by semantic universals, independently of language-specific morphosyntax? the first objective of our study is methodological: to develop a task that directly asks participants about interpretation, targeting the concept of atomicity, and does not provide any morphosyntactic cues to interpretation. the second objective is theoretical: to examine whether there are differences in the interpretation of flexible nouns between english and gc languages, in the absence of morphosyntactic cues. 2. background. in this section, we first discuss how interpretation of count and mass nouns has been experimentally studied in prior literature, and then move into the specifics of the three languages under investigation in our study. 2.1. prior studies on the interpretation of count and mass nouns. prior studies have used two types of tasks to address the relationship between the object/substance distinction and count/mass morphosyntax: the object/substance rating task and the quantity judgment task. barner, inagaki and li (2009) used the object/substance rating task with native speakers of english and of japanese, asking them to rate 100 common nouns with regard to whether they denote objects, substances, both, or neither. all nouns were presented in bare singular forms (no plural marking or determiners). despite clear morphosyntactic differences between english and japanese, the two participant groups performed very similarly: nouns that are count in english (e.g., ball) were classified as denoting objects in both languages; nouns that are mass in english (e.g., water) were classified as substance-denoting in both languages; and many food-denoting nouns (e.g., pizza, banana) received a mix of ratings, in both languages. in the quantity judgment task, used in barner and snedeker (2005) and many subsequent studies, participants are shown two characters, one of whom has two large objects (e.g., two large chairs or two large blobs of mustard), while the other has six small objects (e.g, six small chairs or six small blobs of mustard), and are asked "who has more x?" the choice of two large objects for the answer corresponds to a judgment by volume, while the choice of six small objects corresponds to a judgment by number. this paradigm has been used in english by barner and snedeker (2005), barner, et al. (2009), inagaki and barner (2009), and macdonald and carroll (2018), and in french by inagaki and barner (2009). the paradigm has also been extended to gc languages: it was used in japanese by barner, et al. (2009) and inagaki and barner (2009); in mandarin chinese by cheung, li and barner (2010); and in korean by macdonald and carroll (2018). category sample noun explanation 1. object-count chair atomic nouns, count-cross-linguistically 2. flexible-count bean flexible nouns cross-linguistically, count in english 3. flexible chocolate flexible nouns, both count and mass in english 4. object-mass furniture superordinate nouns, mass in english 5. flexible-mass spinach flexible nouns cross-linguistically, mass in english 6. substance-mass mustard non-atomic nouns, mass cross-linguistically table 2: different noun categories, based on cross-linguistic behavior proceedings of elm 1: 113-124, 2021 sea hee choi and tania ionin: what’s the smallest part of spinach? a new experimental approach to the count/mass distinction. 114 https://doi.org/10.3765/elm https://www.elm-conference.net/ each study investigated some of the noun categories listed in table 2. note that categories 2 through 5 in table 2 are labeled based on their behavior in english, but vary with regard to their count/mass morphosyntax cross-linguistically. in contrast, categories 1 and 6 appear to be universal: clearly bounded objects like chair are count in any language which has a count/mass distinction, while substance-denoting nouns like mustard are always mass.the findings of the studies that have used the quantity judgment task can be summarized as follows. for categories 1, 4 and 6, there is strong evidence of universality: speakers of both plural-marking and gc languages uniformly judge nouns in categories 1 and 4 by number, and those in category 6 by volume. in contrast, there is much cross-linguistic variability with regard to categories 3 and 5 (category 2 has not, to the best of our knowledge, been tested with the quantity judgment task). in plural-marking languages, judgments are dependent on the morphosyntax: e.g., in english, chocolates (count) is judged by number, while spinach and chocolate (mass) are judged by volume, whereas in french spinach is count and is judged by number. in gc languages, the judgments for these noun types fall in-between. thus, the results of the quantity judgment task suggest that there is both universality (nouns which are always object-denoting, such as chair/furniture are always judged by number) and the effects of language-specific morphosyntax. a potential limitation of the quantity judgment task is that, in plural-marking languages such as english and french, the task confounds morphosyntax with interpretation: count nouns are presented in plural form ("who has more chairs / chocolates?") while mass nouns are presented in singular form ("who has more mustard / chocolate / spinach?"). in gc languages, on the other hand, all nouns are presented in bare form, with no plural marking. thus, it is possible that the apparent effect of morphosyntax on interpretation (the finding that spinach is judged by volume in english but by number in french) is due to the task format, specifically, to whether the noun appeared in singular or plural form in the task. in light of this, we have developed a new task, the minimal parts identification task (mpit), which avoids this problem by presenting all nouns in bare singular form. additionally, the mpit aims to probe directly into speakers' judgments of atomicity, by asking participants about whether a given entity has minimal parts. 2.2. the count/mass distinction in english, korean and mandarin. the three languages examined in the present study are english, korean and mandarin chinese. as discussed above, english is a plural-marking language with an obligatory count/mass distinction. as shown in tables 1 and 2, the count/mass distinction only partially corresponds to the semantic atomicity distinction. english has nouns like furniture, which are mass despite denoting atomic entities. english also has nouns like string(s), stone(s), etc., which have both count and mass variants. korean and mandarin are both gc languages, and there is much debate as to they have a grammaticized countmass distinction (see chierchia 1998b, 2010; cheng and sybesma 1998). both languages have plural marking, but the korean plural marker -tul has a much wider distribution than the mandarin plural marker -men. according to kim (2005), korean -tul is directly related to atomicity, being compatible with atomic nouns but not with non-atomic ones. this claim was supported by experimental findings in choi, ionin and zhu (2018), who found that native korean speakers accepted -tul with object-denoting nouns like chair as well as those like furniture (categories 1 and 4 in table 2), but not substance-denoting nouns like oil (category 6). at the same time, -tul is nearly always optional, as discussed by kim (2005) and kwon and zribihertz (2004), among others, so that a bare singular noun is compatible with both singular and plural interpetations (one exception is definite contexts, where -tul is obligatory for plural interpretation). in the case of mandarin, the plural marker -men is restricted to [+human] nouns; for more discussion, see iljic (1994) and li (1999). at the same time, according to cheng and proceedings of elm 1: 113-124, 2021 sea hee choi and tania ionin: what’s the smallest part of spinach? a new experimental approach to the count/mass distinction. 115 https://doi.org/10.3765/elm https://www.elm-conference.net/ sybesma (1998, 1999), the atomicity distinction is reflected in the classifier system, with clear syntactic differences in the behavior of count and mass classifiers. thus, we have three languages with differences both in the distribution of plural marking, and the correspondence between atomicity and morphosyntax. if interpretation is at least partially influenced by morphosyntax, then we would expect speakers of english, korean and mandarin to exhibit differences in their judgments of atomicity. if, on the other hand, interpretation is universal and independent of morphosyntax, then we would expect very similar behavior from all three language groups. before we can study interpretation, however, we need to establish exactly how the different noun types in table 2 behave with regard to morphosyntax. while none of the noun types in table 2 are compatible with -men in mandarin (since they are all [-human]), it is an open question as to which of these noun types are compatible with -tul in korean. choi, et al. (2018) found that -tul was compatible with the nouns in categories 1 and 4, but not category 6; however, they did not test categories 2, 3 and 5. in this study, we administered a grammaticality judgment task (gjt) to native korean speakers in order to examine the compatibility of -tul with all noun types in table 2; to allow for a cross-linguistic comparison, we administered the gjt in english as well. 3. noun selection. in this study, we tested the six noun types in table 2 with regard to both morphosyntax (the gjt, section 4) and interpretation (the mpit, section 5). before we present those tasks, we discuss how the nouns for the six categories in table 2 were selected. for categories 1 and 6, we selected clearly object-denoting and clearly substance-denoting nouns, which are expected to be universally count and mass, respectively. for category 4, we included superordinate object-denoting nouns like furniture, jewelry, etc., which we had investigated in prior studies on the second language acquisition of the count/mass distinction; see choi, et al. (2018) and choi and ionin (in press). for category 3, we selected nouns which have both singular and plural variants in english (string(s), chocolate(s), etc.). the main challenge was finding nouns for categories 2 and 5, which are invariably count and mass, respectively, in english, yet have different status in other languages. we created a list of nouns which have the potential to be flexible cross-linguistically (primarily names of various fruit and vegetables, foods and materials), and asked linguists from eight obligatory plural marking languages other than english (spanish, russian, german, greek, brazilian portuguese, polish, french, basque) to categorize the nouns as either ‘count’, ‘mass,’ or ‘flexible’ in their language. based on the results of this survey, we selected for category 2 those nouns that are count in english but that were categorized as mass or flexible in at least one plural-marking language; for category 5, we selected those nouns that are mass in english but that were categorized as count or flexible in at least one plural-marking language. a total of eight nouns were selected for each noun type in table 2, 48 nouns total. 4. grammaticality judgment task. the goal of the gjt was to establish the behavior of plural marking in both english and korean, in order to relate it to interpretation of different noun types in those languages (see section 5). mandarin was not tested in the gjt, since, as discussed above, the mandarin plural marker -men is incompatible with [-human] nouns. 4.1. participants. 20 native speakers of english residing in the u.s. (mean age = 21), and 20 native speakers of korean residing in south korea (mean age = 22) completed the gjt. about half of the participants in each group had also completed the mpit about a month earlier. the english participants were recruited via amazon's mechanical turk, while the korean participants were recruited via online advertisement. all the participants were tested online via the survey gizmo tool. proceedings of elm 1: 113-124, 2021 sea hee choi and tania ionin: what’s the smallest part of spinach? a new experimental approach to the count/mass distinction. 116 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4.2. materials. prior to the actual test, the sentence frames for the gjt were normed to make sure that they were not biased towards object-denoting or substance-denoting interpretations. the norming was done via mechanical turk, with 15 native english speakers who did not take part in the main experiment. the participants in the norming task were asked to read sentence frames containing x in the object position (e.g., we talked about x at school yesterday) and judge the likelihood of x being an object or a substance on a scale from 1 (definitely a substance) to 5 (definitely an object). the four sentence frames with the most neutral ratings were selected for the gjt. these were: we talked about x at school yesterday; we read about x in the library yesterday; we searched for x online yesterday; and i thought about x when i was in the kitchen yesterday. for the gjt, 48 target sentences were created; each sentence corresponded to one of the four neutral sentence frames, with the target noun in object position, in place of x. there were 48 target nouns, corresponding to the six conditions in table 2, eight nouns per condition. two versions of each sentence were created, one with the singular and the other with the plural form of the noun; no determiners were used with the target nouns. two experimental lists were created, and the singular and plural versions of each sentence were distributed across the two lists using a latinsquare design. each list contained 30 filler items in addition to the 48 target items; the items were pseudo-randomized for order of presentation. a sample item for one of the categories is given in (1a) for english, and (1b) for korean. the plural marker is given in parentheses here; in the actual test, a given sentence either did or did not contain the plural marker. (1) sample item for category 2: a. we talked about bean(s) at school yesterday. b. wuli-nun ecey hakkyo-eyse khong-(tul)-eytayhay iyaki hay-ss-ta. we-top yesterday school-loc bean-pl-dat about talk-pst-decl. the participants were asked to rate the grammaticality of each sentence on a scale from 1 (not acceptable) to 4 (very acceptable). in english, nouns from categories 1 and 2 were expected to be acceptable in plural form and unacceptable in singular form, while the opposite was expected to be the case for nouns in categories 4, 5 and 6; in category 3, both singular and plural forms are grammatical. in korean, the bare singular form is always grammatical; based on the results in choi, et al. (2018), the plural -tul form was expected to be more acceptable in categories 1 and 4 than in category 6; its acceptability in categories 2, 3 and 5 was an open question. 4.3. descriptive results. figure 1 shows the mean ratings of all six noun types in singular and plural forms in both languages. english speakers performed as expected given english morhposyntax, rejecting the singular form in categories 1 and 2, rejecting the plural form in categories 4 through 6, and accepting both forms in category 3 (though the plural was rated higher than the singular in category 3, the singular form still received a mean rating of about 3 on a 1-to4 scale). korean speakers always accepted the bare singular forms, as expected. the plural -tul forms were fully acceptable in categories 1 and 4, less acceptable in categories 2 and 3, and received the lowest ratings in categories 5 and 6. in sum, the english and korean speakers had similar judgments of count/mass morphosyntax with flexible-mass and substance-mass nouns, but differed in the other four categories. we now move on to the question of whether speakers of these languages (as well as of mandarin) also differ on their interpretation of count vs. mass nouns. proceedings of elm 1: 113-124, 2021 sea hee choi and tania ionin: what’s the smallest part of spinach? a new experimental approach to the count/mass distinction. 117 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: gjt results for english and korean: mean ratings 5. minimal parts identification task. the mpit is a new task specially devised to study atomicity. as discussed earlier, the main advantage of the mpit over the quantity judgment task from barner and snedeker (2005) is that the mpit asks about atomicity (minimal parts) more directly and does not provide morphosyntactic cues, thus addressing interpretation rather than morphosyntax. 5.1. participants. 20 native speakers of english in the u.s. (mean age = 21), 20 native speakers of korean in south korea (mean age = 23) and 20 native speakers of mandarin chinese in china (mean age = 20) completed the mpit. the english speakers were recruited via amazon’s mechanical turk; the korean speakers and the mandarin speakers were recruited in online university communities, via advertisement. all participants were tested online via the survey gizmo tool. 5.2. materials. the mpit materials were first created in english, and then translated into korean and mandarin. participants first read the instructions in (2a); for each test item, they responded first to the question in (2b); if they answered yes to this question, they were asked to provide the name of the minimal unit, in response to (2c). only responses to (2b) are reported in this paper; responses to (2c) have not yet been analyzed. the same 48 nouns were used in the mpit as in the gjt (the six categories in table 2, eight nouns per category). the 48 items were pseudorandomized for order of presentation. each noun appeared in singular bare form, as shown in (2b). there was only one experimental list. (2) a. instructions (english version): when we see something, we can sometimes think of its minimal (smallest) unit. for example, the minimal unit of table is a table: if you divide a table in half, it cannot function as a table anymore. however, the minimal (smallest) unit of water is vague (it is unclear what the minimal (smallest) unit is): if we divide water in half, it will still be water. for each item in this task, you will see a word and two questions about the given word. in the first question, please indicate whether you can think of the minimal (smallest) unit for this word, by clicking either ‘yes’ (you can think of the minimal unit, as with table), or ‘no’ (you cannot think of the minimal unit, or the minimal unit is vague, as with water). in the second question, you will see another question which will ask what is the minimal (smallest) unit of the given object/substance. if you answered “yes” to the first question, then, please type what you think the minimal unit is. for proceedings of elm 1: 113-124, 2021 sea hee choi and tania ionin: what’s the smallest part of spinach? a new experimental approach to the count/mass distinction. 118 https://doi.org/10.3765/elm https://www.elm-conference.net/ example, if you see the word “table”, you will click ‘yes’ and type ‘a table’. if you answered “no” to the first question, please type “n/a” “vague” or “none” in the second question. b. does chair have a minimal unit? (yes)/(no) c. if yes, what is the minimal (smallest) unit? 5.3. predictions. as discussed in section 2.1, prior studies with the quantity judgment task have found striking uniformity with regard to nouns in categories 1, 4 and 6. we expect the same universality to be manifested in the mpit: the responses to the question in (2b) should be primarily yes (there is a minimal unit) for categories 1 and 4 and primarily no (there is no minimal unit) for category 6, in all three languages. with regard to the other three categories, there are two possibilities. one possibility is that morphosyntax drives interpretation, as was found in studies using the quantity judgment task. in that case, english speakers should give primarily yes responses to category 2, and primarily no responses to category 5. it is less clear how they would respond to category 3 (flexible nouns like chocolate(s)). since all nouns are presented in bare form in the mpit, participants see chocolate (singular, mass) rather than chocolates (plural, count), and, under the influence of morphosyntax, would therefore be more likely to give a no response. alternatively, english speakers may consider both count and mass interpretations of the noun, and therefore give a mix of yes and no responses. moving on to korean speakers, figure 1 shows that for categories 2, 3 and 5, they rated the singular form higher than the plural, with no obvious differences among the three categories. therefore, if morphosyntax drives interpretation, korean speakers should behave about the same on categories 2, 3 and 5; the same holds for mandarin speakers, for whom all three categories are equally incompatible with the plrual marking -men. exactly what response type might be expected from the korean and mandarin speakers on categories 2, 3 and 5 is an open question, but crucially, they should behave differently than english speakers, who should distinguish among the three categories. alternatively, it is possible that interpretation is universal and largely independent of morphosyntax, and that the cross-linguistic differences obtained on flexible nouns with the quantity judgment task (barner et al. 2009, inagaki and barner 2009) was due to the specific nature of that task, in which count nouns were presented in plural form but mass nouns in singular form in english, while all nouns were presented in bare singular form in gc languages. in that case, we expect to find no cross-linguistic differences for any category in the mpit, where all nouns are presented in bare singular form in all three languages. 5.4. results. figure 2 presents the mpit results, as the percentage of yes (there is a minimal unit) responses to each noun type. overall, the three groups showed very similar performance. at the same time, compared to english speakers, speakers of korean and mandarin showed an overall higher proportion of yes responses to the minimal parts question in almost every condition, with mandarin speakers showing the highest proportion of yes responses among all three groups. proceedings of elm 1: 113-124, 2021 sea hee choi and tania ionin: what’s the smallest part of spinach? a new experimental approach to the count/mass distinction. 119 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2: mpit results (% of yes, there is a minimal unit responses) for the statistical analysis, the dependent variable (the response to (2b)) was coded with 1 for yes and 0 for no. the data were analyzed using a mixed effect logistic regression model in r, see bates, mächler, bolker and walker (2014). the following fixed effects were included in the model: language group (3 levels) and category (6 levels). helmert coding was used for the variable of language group: the english group was compared to the full gc group; subsequently, the korean and mandarin groups were compared to each other. this allowed us to see a comparison between plural-marking and gc languages, as well as between two different gc languages. the variable of category was coded with contrast coding. an interaction between group and category was also introduced in the model. subjects and items were included as random effects. following barr, levy, scheepers and tily (2013), maximal models were created but since the maximal models did not converge the model was gradually reduced; the final model included by-subject and by-item random intercepts. the model output is given in table 3. the performance of the english group differed significantly from that of the gc group, while the korean and mandarin groups did not differ. there were also significant effects of category for most of the category comparisons. the interactions between group and category were significant for several of the english/gc comparisons, and marginal for one of the korean/mandarin comparison. in order to explore the sources of the interactions, bonferroni pairwise comparisons were conducted using the emmeans() function in r, see lenth (2020) (the bonferroni correction for multiple comparisons is automatically implemented in r) as well as a visual examination of the interaction plots using the plot_model() function, see gelman (2008). with regard to cross-group comparisons, differences were noted only in category 4 (object-mass nouns), where mandarin speakers gave significantly more yes responses than english speakers, and marginally more yes responses than korean speakers; and in category 5 (flexible-mass nouns), where mandarin speakers gave marginally more yes responses than english speakers. there were no group differences on any of the other categories. 1.count 2.fl-count 3.flexible 4.obj-mass 5. fl-mass 6. mass chair bean chocolate furniture spinach mustard proceedings of elm 1: 113-124, 2021 sea hee choi and tania ionin: what’s the smallest part of spinach? a new experimental approach to the count/mass distinction. 120 https://doi.org/10.3765/elm https://www.elm-conference.net/ fixed effects coefficie se z p-value intercept .41 .17 2.37 .01* group (english vs. gc) 1.14 .41 2.81 <.01* group (korean vs. mandarin) .35 .35 1.01 .31 category (1 vs. 2-6) 2.48 .16 15.81 <.01* category (2 vs. 3-6) .92 .11 8.53 <.01* category (3 vs. 4-6) -.41 .10 -4.19 <.01* category (4 vs. 5-6) -.13 .10 -1.31 .19 category (5 vs. 6) -.05 .10 -.47 .64 group (english vs. gc) x category (1 vs. 2-6) .48 .40 1.20 .23 group (korean vs. mandarin) x category (1 vs. 2-6) -.66 .42 -1.58 .11 group (english vs. gc) x category (2 vs. 3-6) -.61 .30 -2.05 .04* group (korean vs. mandarin) x category (2 vs. 3-5) -.10 .27 -.36 .72 group (english vs. gc) x category (3 vs. 4-6) -.69 .27 -2.53 .01* group (korean vs. mandarin) x category (3 vs. 4-6) .35 .24 1.45 .15 group (english vs. gc) x category (4 vs. 5-6) .26 .27 .99 .32 group (korean vs. mandarin) x category (4 vs. 5-6) .41 .24 1.75 .08+ group (english vs. gc) x category (5 vs. 6) -.13 .28 -.48 .63 group (korean vs. mandarin) x category (5 vs. 6) -.21 .24 -.88 .38 random effects variance sd subject 1.07 1.03 item .08 .29 * significant at p<.05 + marginal at .05 accuracy of uqs in the two groups. rq2) is there an interaction between np type and age? is the joint effect of np type and age on accuracy predictable from the effect of the two factors individually? – h0: there is no interaction. – h1: there is an interaction. 3.2. method. 3.2.1. participants. a total of 55 spanish-speaking children (30 male; 25 female) divided into two age groups, a 4/5-year-old group (n = 31, m = 68.16 months, sd =6.8, henceforth “young”) and an 8/9-year old group (n = 24, m = 108.75 months, sd = 6.3, henceforth “old”), were recruited from a local (primary) school in vitoria-gasteiz (spain). the participants in the young group belonged to different school years: 2nd and 3rd years of “infantil” (child-care). alongside, we included a third group of adult controls (n = 26, 12 male; 14 female, henceforth “adults”). we did not test college undergraduates, but rather adults of various ages (m = 38, sd = 13.92) and in various schooling degrees (ranging from elementary education to a ba university degree). all participants were spanish speakers and residents in vitoria-gasteiz. all volunteered to take part in the experiment. this study was carried out in accordance with the recommendations of the human beings research ethics committee (“ceish: comité de ética de investigación con seres humanos de la upv/ehu”). 3.2.2. design. building on lazaridou-chatzigoga & stockall (2013) and lazaridou-chatzigoga et al. (2019), the design we proposed placed the role of exceptions at the center of the discussion. to keep the experiment simple, we also focused only on majority characteristic statements. we manipulated np into a generic (det.pl) and a universal (todos/as det.pl ‘all’) condition. by contrast, instead of manipulating the context as in lazaridou-chatzigoga et al. (2019), we did not manipulate the context variable between a supportive and a contradictory condition; all our critical items were uttered on the basis of a contradictory context, i.e., in the face of exceptions. below is an example. (5) critical condition a. context: participant sees the picture of a rabbit with one ear and hears the utterance “a rabbit with one ear”. proceedings of elm 1: 090-100, 2021 elena castroviejo, dimitra lazaridou-chatzigoga, marta ponciano and agustı́n vicente: generics as default? comparing the acquisition of universals and generics in spanish. 93 https://doi.org/10.3765/elm https://www.elm-conference.net/ b. target question: ¿dirı́as que {todos los, los} conejos tienen dos orejas? say.cond.2sg that all det.pl det.pl rabbits have.3pl two ears ‘would you say that {all, ∅} rabbits have two ears?’ however, to prevent participants from establishing a pattern, we included supportive contexts in the fillers. each participant in the study saw 16 critical items, 32 distractors. in the critical items, np type (gs vs. uqs) was a within-individual independent variable. we also had a between-individual variable, namely the three age conditions: young, old and adult. the dependent variable was a yes/no response to the question prompted as part of the experimental procedure. given the manifest existence of exceptions and following the semantic-logical properties of the different expressions, gs and uqs were true in different scenarios. for instance, given the context described in (5) and the learned truth-conditions, the accurate answers are the ones in (6). (6) a. all rabbits have two ears. → no b. rabbits have two ears. → yes coming back to the distractors, we created two conditions: (a) similar sentences with proper name subjects (henceforth “name”), (7), which was a mix of true and false statements, to control for the participants’ yes-bias and for their understanding of the procedure, and (b) generic-like statements that adults tend to reject in generic form even though they predicate prevalent properties of a kind. these statements have been called “false generalisations” by leslie et al. (2011). we used them as control generalisations with a supportive context (henceforth “controlgen”), (8). (7) name condition a. context: participant sees a picture of plaza de la virgen blanca and hears a voice uttering “plaza de la virgen blanca”. b. target question: ¿dirı́as que la plaza de la virgen blanca está en vitoria? say.cond.2sg that det.sg plaza de la virgen blanca is in vitoria ‘would you say plaza de la virgen blanca is in vitoria?’ (8) controlgen condition a. context: participant sees a picture of a black roof and hears a voice uttering “a black roof”. b. target question: ¿dirı́as que los tejados son negros? say.cond.2sg that det.pl roofs are black.pl ‘would you say roofs are black?’ 3.2.3. materials. as previously mentioned, participants saw 16 critical items (8 gs and 8 uqs) and 32 distractors (16 names and 16 controlgen). they were distributed in two lists that were randomly assigned to the participants. so, each critical item came in two different conditions (gs or uqs), and participants in the different lists did not see the same item in the same condition. all participants saw the same set of distractors. apart from the set of 16 critical items proceedings of elm 1: 090-100, 2021 elena castroviejo, dimitra lazaridou-chatzigoga, marta ponciano and agustı́n vicente: generics as default? comparing the acquisition of universals and generics in spanish. 94 https://doi.org/10.3765/elm https://www.elm-conference.net/ and 32 distractors, there were 4 training items, also common to all participants. the materials were counterbalanced across participants. also, experimental items were randomized every time a participant started a new experimental session. the 16 critical items consisted of majority characteristic statements like cats have whiskers and horses have four legs. special care was taken to select properties about which young children would be knowledgeable. also, to avoid the possible ambiguity in the gs condition (det.pl can be ambiguous between a generic and an exemplar reading in spanish), the picture that was shown always depicted a single individual holding an exceptional property. also, bear in mind that, for each critical generalization, a picture describing an exception was presented. to create these materials we used two strategies: we either selected a subkind of the species (a sphinx cat, which does not have whiskers, or a black pig, which is not pink) or an animal that, for an accidental reason, may have suffered a mutation (a three-legged horse or a hen with four wings). regarding controlgen, unlike gs, they described characteristics held of only some of the instances of a kind. for instance, square-shaped pizzas, black roofs or blue butterflies. in these cases, accompanying pictures were supportive, so the participant would see a picture of a squareshaped pizza while she was asked whether pizzas are square-shaped. 3.2.4. procedure. children were tested individually in a quiet room in their school. they had been previously told that they would play a game on the computer. in the case of adults, they were administered the study at a quiet place of their convenience. participants sat in front of a computer screen with the investigator beside them. in the case of children, the experimenter was in charge of clicking the left or right key on the computer’s mouse, corresponding to a “yes” or “no” answer, while adults handled the mouse themselves. the software used was e-prime 3 (“psychology software tools”) on a pc running windows. no feedback was given during the main task. the testing process took approximately 15 minutes to complete. in the study, each target item was composed of two parts. in the first part, an image occupying the center of the screen presented one individual contradicting or supporting the generalization to be judged. for instance, a cat without whiskers. at the same time, a pre-recorded audio of a female voice said: “un gato sin bigotes” (‘a cat without whiskers’). by pressing a key, we moved to the second part, where an image of a girl, a cartoon character, appeared on the right-hand-side of the screen and asked “¿dirı́as que (todos) los gatos tienen bigotes?” (‘would you say that (all) cats have whiskers?’). the pre-recorded audio with the question was played twice. it was a twoalternative forced choice task, whereby participants were instructed to choose “yes” or “no” in light of the picture that had been presented to them. 3.2.5. data analysis. we used the statistic tools from ibm spss 26 to analyze the variance of the means of accurate responses per item by running repeated measure anovas with np type as a within-individuals factor and age as a between-individuals factor. 3.3. results. table 1 summarizes the mean accuracy and standard deviation of responses by np type (gs vs. uqs) and age group (young, old, adult) in a by-item analysis. figure 1 graphically illustrates the critical differences. proceedings of elm 1: 090-100, 2021 elena castroviejo, dimitra lazaridou-chatzigoga, marta ponciano and agustı́n vicente: generics as default? comparing the acquisition of universals and generics in spanish. 95 https://doi.org/10.3765/elm https://www.elm-conference.net/ gs uqs young 0.92 [sd 0.08] 0.30 [sd 0.18] old 0.72 [sd 0.08] 0.64 [sd 0.25] adults 0.79 [sd 0.19] 0.72 [sd 0.20] table 1: descriptive statistics (critical) figure 1: mean accuracy per np type and age group we conducted a repeated measure anova on the control categories, crossing np type (gs vs. uqs), and age (young, old, adult). the analysis revealed a significant effect of age (f(2.45) = 8.057, p = 0.001, η2 = 0.264) and a significant effect of np type (f(1.45) = 35.709, p < 0.001, η2 = 0.442). further, the analysis yielded a significant interaction between age and np type (f(2.45) = 17.357, p < 0.001, η2 = 0.435). the magnitude of the age effect is small, but the effect size of np type and the interaction is close to medium. to be able to properly interpret this interaction, we performed pairwise comparisons. the pairwise comparisons between the two levels of the factor np type (gs vs. uqs) corrected for bonferroni, yielded the result that the accuracy of gs is significantly higher than the accuracy of uqs. if we compare the age factor within the two levels of the np type factor, we observe that only in the young group there is a significant difference between gs and uqs, such that the accuracy of gs is significantly higher than uqs. additionally, accuracy in gs differs significantly between the young group on the one hand, and the old group (p = 0.000) and adults (p = 0.025) on the other hand, in favor of the young group. finally, with respect to uqs, there is a significant difference (p = 0.000) between the young group on the one hand, and old group and adults on the other hand, this time in favor of the latter groups. to obtain results from the pairwise comparisons within the between-participants factor (namely, age), we ran the games-howell test (since we could not assume equality of variance), and the results yielded that accuracy differs significantly only between the young and adult conditions (p = 0.000), but not between young and old or between old and adult. moving to the distractor items, table 2 presents the mean accuracy and standard deviations of the three age groups depending on filler type. this is graphically represented in the bar graph in figure 2. from the results in the name condition, it is obvious that all participants understood the task, were paying attention and had some amount of world knowledge (it increases with age). more revealing is the controlgen condition, which shows very low measures in the children groups, proceedings of elm 1: 090-100, 2021 elena castroviejo, dimitra lazaridou-chatzigoga, marta ponciano and agustı́n vicente: generics as default? comparing the acquisition of universals and generics in spanish. 96 https://doi.org/10.3765/elm https://www.elm-conference.net/ whereas a comparatively high value in the case of adults. controlgen name young 0.49 [sd 0.20] 0.81 [sd 0.10] old 0.44 [sd 0.14] 0.92 [sd 0.09] adults 0.84 [sd 0.10] 0.97 [sd 0.04] table 2: descriptive statistics (distractors) figure 2: mean accuracy per np type and age group (distractors) even though gs were counterbalanced in two different lists while the list of controlgen was common for all participants, we decided to conduct a repeated measure anova to compare gs and controlgen across the age factor and thus explore the significance of differences, even if only as a tentative measure. the analysis revealed a significant interaction between type of generic (“gentype”) and age. the analysis revealed significant differences of gentype (f(1.45) = 47.736, η2 = 0.515), age (f(2.45) = 22.802, η2 = 0.503) and their interaction (f(2.45) = 19.822, η2 = 0.468). regarding pairwise comparisons within gentype, the young and old group differ significantly in their accuracy of gs vs. controlgen (p = 0.000), but not the adult group. with respect to the age factor, there were significant differences in all of them. the difference between young and old group yielded a p = 0.004, between young children and adults, the comparison yielded a p = 0.013, and the difference between older children and adults had a p = 0.000. 3.4. discussion. let us start by addressing the research questions that we spelled out in 3.1. rq1) are children sensitive to the reported differences between gs and uqs? rq2) is there an interaction between np type and age? is the joint effect of np type and age on accuracy predictable from the effect of the two factors individually? in both cases, the null hypothesis can be rejected. first, considering all age groups, the accuracy in gs is greater than the accuracy in uqs. second, there is an interaction between age and np type such that the difference in accuracy between gs and uqs is much larger in the case of young children than old children and adults. now, do these data support the gad view? one of its main claims is that generics are held to be true even if we are aware that there are exceptions. as we have mentioned above, this may seem supported by the data at first sight. however, we have observed that generics are not always verified as they should, so this is not true across development. even in the case of majority proceedings of elm 1: 090-100, 2021 elena castroviejo, dimitra lazaridou-chatzigoga, marta ponciano and agustı́n vicente: generics as default? comparing the acquisition of universals and generics in spanish. 97 https://doi.org/10.3765/elm https://www.elm-conference.net/ characteristic generics, the ones we have tested, generics are not universally taken to be true. in fact, some of our reported data seem to go against it. especially, the high success of the young group in the gs condition (in fact higher than the old group and adults) shows there is a significant decline in the accuracy of gs between the young and old groups and even extending to adults. if the greater accuracy of gs in young children, both with respect to the uqs condition and other ages, were to be interpreted as evidence in favor of gs being easy, we would be forced to entertain the idea that gs become more difficult across development, which is something that we would not want to argue for. moreover, the low rates in the controlgen condition in the children groups clearly show that they behave unlike adults, so there seems to be a development trajectory in the adult-like understanding of gs, which is not predicted under the hypothesis that verifying a gs involves system 1 and, as such, the correct interpretation of gs should be preserved across ages. in fact, given the poor performance of young children in the uqs condition, the old group is adultlike in their interpretation of uqs, while the output in the generic types taken together (gs and controlgen) suggest the two children groups are non-adult-like in the interpretation of generics. finally, the mere fact that the old group and adults do not have at ceiling results for gs can also be viewed as data that go against the gad (i.e., if gs are easy, we would expect higher accuracy than the one we observe). on the other hand, it should be noted that the centrality of exceptions and the fact that the participants had to reason in the face of a counterexample might have influenced these non-at-ceiling effects. remember that gad is an answer to the generic overgeneralization (gog) effect that was observed in various experiments. now, do we also find a gog effect in our data? the low values in uqs in the young group could be evidence in favor of this, since it seems they are interpreting them as tolerating exceptions. the gog effect would be attenuated in the older group and in adults, but the low performance in uqs in these two groups may also suggest that they tend to interpret some uqs as gs. as a final remark, we may be tempted to make sense of this developmental cline with respect to gs as describing a u-shape (as has been claimed e.g. by berko (1958) for the acquisition of english past tense morphology); that is, one could imagine that young children easily acquire gs and then, when they become proficient with uqs, they start having doubts about the relevance of exceptions in generalizations, so they start rejecting gs when they should not. we believe the collected data suggest otherwise, especially if we take into account the results of controlgen and the lack of ceiling effects in adults. in fact, the at-chance rate of acceptance observed in the young group with controlgen (remember that controlgen includes items that the literature on generics would consider not to be acceptable as generic generalizations) seem incompatible with the idea that young children have an adult-like command of generics altogether. we also can’t interpret our results as providing evidence that the old group behaves differently from young and adults taken together. the old group and adults pattern together in gs, uqs and name, while the young and old groups pattern together in controlgen. since even in the odd case of a u-shaped curve the gad theory would not be able to account for the data we present, we need to find other theories that are compatible with them. 4. conclusions. in the present paper we have carried out an investigation about generalizations in spanish-speaking children of two age groups (and a corresponding group of adult controls). proceedings of elm 1: 090-100, 2021 elena castroviejo, dimitra lazaridou-chatzigoga, marta ponciano and agustı́n vicente: generics as default? comparing the acquisition of universals and generics in spanish. 98 https://doi.org/10.3765/elm https://www.elm-conference.net/ building on previous work by lazaridou-chatzigoga & stockall (2013), lazaridou-chatzigoga et al. (2019), which addressed leslie and gelman’s “generics as default” (gad) hypothesis, we have proposed a design that could test differences between generic statements (realized as definite plurals in spanish) and unrestricted universal quantified statements, when the participants were faced with a photograph of an individual failing to support the generalization. the data that we have collected does not talk in favor of the gad view. in fact, it describes an interesting picture yet to be fully understood. we would like to emphasize that, while it was established that 4-year-olds were able to comprehend (restricted) uqs in an adult-like manner, we have found out that they are not adult-like in the interpretation of unrestricted uqs. in fact, by comparing three age groups, we have been able to spot an age group, namely 8/9-year-olds, as having certain adult-like behaviors (for instance in the interpretation of uqs) and child-like behaviors (for instance in the interpretation of the “false” generics). finally, in this study we have analyzed behavioral results from a forced-choice task. however, in view of the apparent mismatches observed in the literature between the information collected from behavioral tasks and e.g. reaction times, we believe that processing data should be key in further informing us on whether the status of the gad hypothesis. references barberán-recalde, tania. 2019. the acquisition of basque and spanish quantifiers: an empirical study: university of the basque country dissertation. berko, jean. 1958. the child’s learning of english morphology. word 14. 150–177. 10.1080/00437956.1958.11659661. dahl, östen. 1995. the marking of the episodic/generic distinction in tense-aspect systems. in greg carlson & francis jeffry pelletier (eds.), the generic book, 412–425. chicago: chicago university press. gelman, susan a. 2010. generics as a window onto young childrenâs concepts. in francis jeffry pelletier (ed.), kinds, things, and stuff: mass terms and generics, 100–120. new york: oxford university press new york. 10.1093/acprof:oso/9780195382891.003.0006. gelman, susan a., ingrid sanchez tapia & sarah-jane leslie. 2016. memory for generic and quantified sentences in spanish-speaking children and adults. journal of child language 43(6). 1231–1244. https://doi.org/10.1017/s0305000915000483. kahneman, daniel & shane frederick. 2002. representativeness revisited: attribute substitution in intuitive judgment. in t. gilovich, d. griffin & d. kahneman (eds.), heuristics and biases: the psychology of intuitive judgment, 49–81. cambridge university press. https://doi.org/10.1017/cbo9780511808098.004. katsos, napoleon, chris cummins, maria-josé ezeizabarrena, anna gavarró, jelena kuvač kraljević, gordana hrzica, kleanthes k grohmann, athina skordi, kristine jensen de lópez, lone sundahl et al. 2016. cross-linguistic patterns in the acquisition of quantifiers. in proceedings of the national academy of sciences of the united states of america, vol. 113 33, 9244–9249. https://www.jstor.org/stable/26471413. lazaridou-chatzigoga, dimitra, napoleon katsos & linnaea stockall. 2019. contextualising generic and universal generalisations: quantifier domain restriction and the generic overproceedings of elm 1: 090-100, 2021 elena castroviejo, dimitra lazaridou-chatzigoga, marta ponciano and agustı́n vicente: generics as default? comparing the acquisition of universals and generics in spanish. 99 https://doi.org/10.3765/elm https://www.elm-conference.net/ generalisation effect. journal of semantics 36. 617–664. https://doi.org/10.1093/jos/ffz009. lazaridou-chatzigoga, dimitra & linnaea stockall. 2013. genericity, exceptions and domain restriction: experimental evidence from comparison with universals. in emmanuel chemla, vincent homer & grégoire winterstein (eds.), proceedings of sinn und bedeutung 17, 325–343. école normale supérieure. https://semanticsarchive.net/sub2012/lazaridouchatzigogastockall.pdf. leslie, sarah-jane. 2007. generics and the structure of the mind. philosophical perspectives 21. 375–403. https://doi.org/10.1111/j.1520-8583.2007.00138.x. leslie, sarah-jane. 2008. generics: cognition and acquisition. philosophical review 117(1). 1–47. https://doi.org/10.1215/00318108-2007-023. leslie, sarah-jane, sangeet khemlani & sam glucksberg. 2011. all ducks lay eggs: the generic overgeneralization effects. journal of memory and language 65(1). 15–31. https://doi.org/10.1016/j.jml.2010.12.005. lewis, david. 1975. adverbs of quantification. in ed keenan (ed.), formal semantics in natural languages, 3–15. cambridge university press. https://doi.org/10.1017/cbo9780511897696.003. serratrice, ludovica, antonella sorace, francesca filiaci & michela baldo. 2009. bilingual children’s sensitivity to specificity and genericity: evidence from metalinguistic awareness. bilingualism: language and cognition 12(2). 239–257. https://doi.org/10.1017/s1366728909004027. proceedings of elm 1: 090-100, 2021 elena castroviejo, dimitra lazaridou-chatzigoga, marta ponciano and agustı́n vicente: generics as default? comparing the acquisition of universals and generics in spanish. 100 https://doi.org/10.3765/elm https://www.elm-conference.net/ where truth and optimality part. experiments on implicatures with epistemic adverbs adina camelia bleotu, anton benz & nicole gotzner* abstract. in the current paper, we employ a novel shadow play paradigm in order to test romanian monolingual adults’ sensitivity to truth and informativeness and investigate their ability to derive implicatures with epistemic adverbs. we show that implicature rates with epistemic adverbs are higher when participants are asked to reward characters depending on the truth of their statements rather than on whether what they say is the best description of the situation. given participants’ tasksensitivity, we recommend that instructions use optimality criteria, as they are a more sensitive method of probing into implicature generation. keywords. scalar implicatures; modality; epistemic adverbs; truth value judgment task; optimality judgment task; romanian; methodology 1. introduction. experimental paradigms that study pragmatic reasoning by either testing the fit of sentences to situations or the fit of situations to sentences can be divided into two broad groups: they either ask subjects to make judgements about truth and falsity, or they ask them to make judgements about some measure of appropriateness. we argue that paradigms employing judgements about truth and falsity activate reasoning about semantic meaning, while judgments about appropriateness activate reasoning about pragmatic meaning. we consider this hypothesis in the context of a case study that investigates the derivation of scalar implicatures with the epistemic adverb poate ‘maybe’ in romanian in the case of romanian monolingual adults by means of a novel shadow play paradigm. we implement two reward versions of this paradigm, one that asks subjects to reward characters based on judgements about truth and falsity (rightwrong task) and one that asks them to reward characters based on judgements about the appropriateness of sentences (best description task). our experiments show that implicature rates with the epistemic adverb poate ‘maybe’ are significantly higher in the case of the optimality judgment task (best description task) than in the truth value judgment task (right-wrong task), thus emphasizing an important methodological point: that results and, consequently, the theory accounting for them are largely dependent upon the methods used. the paper is organized as follows: after a brief introduction, in section 2, we present previous research on scalar implicatures from a methodological perspective. section 3 deals with previous research on epistemic modality in language acquisition. section 4 describes the experiments we conducted (goals, participants, methodology, results). section 5 discusses the results and their consequences for future research on scales. section 6 draws a conclusion on this basis. * this research was supported by an xprag.de internship offered to adina camelia bleotu at zas berlin within the deutsche forschungsgemeinschaft (dfg) project si games i: experimental game theory and scalar implicatures led by dr. anton benz (grant nr.: be 4348/4-1). anton benz was supported by the bundesministerium für bildung und forschung (bmbf), grant nr. 01ug1411. nicole gotzner was supported by the dfg, grant nr. be 4348/4-2, through the priority program new pragmatic theories based on experimental evidence (spp 1727), and she is further supported by the dfg through the emmy noether programme (grant nr. go 3378/1-1). we are grateful to the students from the faculty of foreign languages and literatures, university of bucharest, who took part in the experiments. authors: adina camelia bleotu, icub, university of bucharest (cameliableotu@gmail.com), anton benz, zas berlin (benz@leibniz-zas.de), & nicole gotzner, zas berlin (gotzner@leibniz-zas.de) and university of potsdam. proceedings of elm 1: 047-058, 2021 c©2021 adina camelia bleotu, anton benz and nicole gotzner published by the lsa with permission of the author(s) under a cc by license. 47 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2. previous research on scalar implicatures. a methodological perspective. an important methodological decision is how to find out whether participants derive implicatures or not (noveck 2001, papafragou & musolino 2003, geurts & pouscolous 2009, clifton & dube 2010, katsos & bishop 2011, van tiel 2013, benz & gotzner 2014, a.o.). experiments are essentially classified into whether they test the fit of sentences to situations (the most frequent method), the fit of situations to sentences, or whether they resort to some indirect measure of the responses given by participants (such as rewards in the best response paradigm in gotzner & benz 2018). given that the literature on implicatures is extremely vast, and covering it is beyond the scope of this paper, our presentation will mainly focus on some relevant examples of the most commonly used paradigm, the paradigm which tests the fit of sentences to situations. in order to find out whether subjects derive implicatures when confronted with a statement like some dogs are black/the dogs may be behind the curtain in a context where all dogs are black/behind the curtain, experimental paradigms initially employed questions used in the original truth value judgment task (crain & mckee 1985, crain & thornton 1998) such as is the puppet right? or do you agree with the statement? (noveck 2001). however, papafragou & musolino (2003:264) challenged this methodology, arguing that such questions actually tap into speakers’ sensitivity to truth: “in our version, instead of asking subjects if the puppet is ‘right’ or ‘wrong’ (as in the original tvjt), we asked whether the puppet ‘answered well’ (i.e., apantise kala, ‘did(she)-answer well?’). this modification was made since we were interested in felicity, not truth.” while running a truth value judgment task calls for questions about right/wrong or true/false, agree/disagree, running a felicity judgment task calls for questions about adequacy/appropriateness. this idea has been explored in various ways in the literature dealing with scalar implicatures, which consequently experienced a methodological shift from truthoriented to felicity-oriented tasks (katsos & bishop 2011). in a sense, even asking if a puppet answered well may be too weak, given that subjects may assess certain underinformative statements as good enough, and, thus, still evaluate matters in terms of truth value. for this reason, experimental pragmatics has been trying to employ novel methods which make subjects sensitive to the difference between optimal statements (the best, most felicitous statements in a certain pragmatic context) and statements that are true but less optimal. among the paradigms testing the fit of sentences to situations, one way of testing optimality rather than truth is by asking subjects to provide graded judgments. katsos & bishop (2011) tested 6to 7-year-old english-speaking children for underinformativeness with existential quantifiers by means of a ternary reward task where children were asked to offer a ‘small’, ‘big’, or ‘huge’ strawberry as a reward to mr. caveman depending on how good the speaker’s responses were. the paradigm revealed sensitivity to underinformativeness on the part of children, who rewarded such statements with big strawberries instead of huge or small ones. given children’s general acceptance of underinformative statements in a standard truth value judgment task (katsos & bishop 2011), the results from the ternary task were quite surprising, revealing that children had more pragmatic sensitivity than previously thought. this led katsos & bishop (2011) to argue that the yes-no binary task was not fine-grained enough to capture the difference between truth and informativeness, and, thus, it gave the illusion that children were insensitive to violations of informativeness, when, in fact, they were merely tolerant. a different manner of implementing graded judgments is by asking participants to rate certain sentences as descriptions of certain proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 48 https://doi.org/10.3765/elm https://www.elm-conference.net/ images by placing the cursor between yes and no. using the cursor placement method, chemla & spector (2011) conducted several experiments testing whether adults derive local implicatures in sentences like every letter is connected to some of its circles (‘every letter is connected to some, not all of its circles.’) and concluded that local implicatures are attested, against geurts & pouscoulous (2009). another way of testing pragmatic adequacy/optimality is the felicity judgment task, which also tests the fit of sentences to situations, but through a forced choice task where participants choose the best sentence out of two sentences. foppolo, guasti & chierchia (2012) conducted a felicity judgment task (experiment 5 in the paper) on 5-year-old italian children in order to investigate their ability to generate implicatures with existential quantifiers. the set-up involved two puppets who were always quarrelling about which of them described various pictures best, and children had to be the judge. for instance, puppet 1 would utter a weak, underinformative statement with qualche ‘some’, while puppet 2 would utter a stronger, informative statement with tutti ‘all’. then, children were asked: “which puppet said it better?”. a similar methodology was used by ozturk & papafragou (2015) in order to test children’s ability to derive implicatures with epistemic may: children had to choose between the statements produced by minnie and donald about the location of an animal in one of two boxes (an underinformative statement with may and a fully informative statement with have to). importantly, the tasks reveal adult-like answers on the part of children, showing that access to stronger alternatives eases pragmatic understanding. there are also paradigms that do not ask for judgments about fit of sentences to situations, or situations to sentences, such as the best response paradigm proposed by gotzner and benz (2018). an important point about the paradigm is that it avoids meta-linguistic judgments about either truth or appropriateness. instead, judgments are read off from the rewards given by participants in an interactive game–theoretic reward task set-up which satisfies grice’s conversational requirements for implicature generation (a recognizable purpose for the talk exchange). in a task where four girls have lost all/some/none of their marbles and they have to find them again, adults have to handle an explicit decision problem, rewarding characters based on statements about how many marbles they find: (i) chocolate for finding all of the marbles, (ii) candy for fewer than all and (iii) a gummy bear for none of the marbles (as a consolation prize). such a task led to many local implicatures in utterances like all of the girls found some of their marbles (‘all of the girls found some, not all of their marbles’), revealing subjects’ sensitivity to subtle considerations of informativeness. this paradigm revealed quite different results from binary truth value judgment tasks (see also benz & gotzner, in press, for an interactive version of the best response paradigm). apart from the notion of felicity, another notion that has been brought under discussion in relation to the adequacy of an underinformative utterance to a pragmatic context is typicality, that is, the degree to which the situation described by the utterance is a typical one (van tiel 2013). typicality has been studied experimentally starting with rosch (1975), who asked participants to produce typicality orderings by evaluating hyponyms of bird (i.e., robin) along a 7-point likertscale. typical members are learnt earlier, recognized faster and more accurately, and produced earlier (van tiel 2013). interestingly, many experiments on implicatures (geurts & pouscoulous 2009, clifton & dube 2010, chemla & spector 2011) are quite similar to rosch’s rating task: participants see a category with several instances/a sentence with several situations and have to decide how well the category/the sentence describes them. for this reason, van tiel (2013) argues that typicality differences may influence the interpretation not only of predicates like bird but also of a quantifier like some. for instance, begg (1987) found that the typical meaning of some is less than half, and degen & tanenhaus (2011) obtained similar results. in the context of associating proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 49 https://doi.org/10.3765/elm https://www.elm-conference.net/ pictures representing various situations with statements, typicality may overlap with felicity or optimality, given that the best statement describing a situation is probably the most typical one as wellthough various aspects seem to impact typicality (e.g. plausibility, world knowledge, a.o.).we note that for our purposes, in the case of epistemic items, both possibility and certainty items are licensed in contexts where one does not have direct access to the objects/animals under discussion. however, using possible instead of certain in a context of certainty would count as not typical, as well as infelicitous and not optimal, given that possible is expected only in uncertainty contexts. the previously noted advantages of methods relying on optimality/felicity lie at the basis of the current experiments on implicatures with epistemic adverbs, contrasting truth sensitivity to pragmatic sensitivity. 3. previous research on epistemic modality. methodological insights from language acquisition. many experimental studies on epistemic modality (hirst & weil 1982, noveck, ho & sera 1996, noveck 2001, ozturk & papafragou 2015 a.o.) have focused on the acquisition of epistemic modal verbs. such studies have shown that children are sensitive to the relative strength of modal verbs from very early on. although aware of the existence of a modal scale, children still have difficulties with modals at age 5, achieving epistemic maturity only later on, around age 7. the paradigm used in the studies on epistemic modality is some version of the hidden object task, where, based on evidence, subjects have to infer the location of a certain hidden object. a first version of the task is represented by the look for the peanut task (hirst & weil 1982), where children were asked to look for a peanut on the basis of certain statements with modals they heard. noveck, ho & sera (1996) and noveck (2001) then tested epistemic items by employing the box paradigm, where objects are hidden in boxes, a paradigm which was later on simplified by ozturk & papafragou (2015). in the initial box paradigm, there were three boxes (two uncovered boxes which contained an animal or two and a covered one), and subjects were interrogated about a third covered box based on disjunctive statements of the type “this box has the same content as either box a or as box b”. ozturk & papafragou (2015) reduced the complexity of the paradigm by resorting to only two boxes and one single animal and by resorting to non-disjunctive input. the semantic forced choice task conducted by noveck, ho & sera (1996) and the semantic truth value judgment task conducted by ozturk & papafragou (2015) both showed that young children have a tendency to reduce uncertainty and accept situations where a stronger statement (with has to) is made instead of a weaker one (with may). in terms of pragmatics, noveck (2001) conducted a truth value judgment task, where subjects had to say whether they agreed with a certain statement, while ozturk & papafragou (2015) ran a felicity judgment task, where subjects had to choose between two statements (an underinformative statement versus a fully informative statement). 5-year-olds performed more adult-like in the felicity judgment task than in the truth value judgment task, which further reinforces the idea that implicatures are more easily accessed by tasks focusing on pragmatic adequacy. interestingly, regardless of the task type, the implicature rates with modals were quite high for adults (close to 90%). nevertheless, in an adaptation of noveck (2001) on epistemic adverbs in romanian (bleotu 2019), many adults were too cautious, rejecting statements about the certainty of something they could not see. 4. current experiments: truth and optimality in the shadow play paradigm. 4.1. rationale and goals. given subjects’ caution in the hidden object paradigm, i.e., in situations of no direct access to the object, we developed a novel shadow play paradigm, where subjects have to reward a dragon for the statements he makes about the identity of a shadow/silhouette, on the basis of certain evidence. importantly, unlike in the hidden object proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 50 https://doi.org/10.3765/elm https://www.elm-conference.net/ paradigm, in the shadow play paradigm, subjects can infer that the shadow must belong to an animal by looking at its silhouette and also because of sounds that accompany it (e.g., woof-woof, representing a dog). by making indirect evidence more direct and, thus, making the task more evidential, we aimed to prevent subjects from being overly cautious and giving unreliable answers. the paradigm itself draws inspiration from shadow play theater, an ancient art form with silhouettes. a similar but somewhat simpler paradigm was also employed by heizmann (2006) in order to test whether children make indirect inferences with necessity modals in english and german. heizmann (2006) used real characters and their silhouettes in order to see how children interpret questions such as who must be eating a banana? versus who must eat a banana?. the reward task was inspired from katsos & bishop (2011) but, instead of using a ternary reward system, we used a binary reward system associated with different linguistic input. in order to test whether subjects are more sensitive to underinformativeness than to truth value in deriving scalar implicatures, we decided to run the same shadow play test in two different versions: (1) a right-wrong task, where subjects were asked to reward a baby dragon with a big/small apple depending on the truth value of his statement, and (2) a best description task, where subjects had to reward a baby dragon with a big/small apple depending on whether what he said was the best description of the situation or not. the expectation was that participants would derive more implicatures in the best description task, which encourages pragmatic sensitivity. 4.2. methodology. 4.2.1. participants. the right-wrong task was conducted on 64 native romanian speakers, and the optimality test was conducted on 63 romanian native speakers, recruited from 1st and 2nd year students at the faculty of foreign languages, university of bucharest. 4.2.2. materials. we implemented two versions of our experiment in penncontroller (zehr & schwarz 2018). while the experiments employ the same type of task (a reward task), the criteria for rewarding were different: truth value (“right-wrong”) (in the right-wrong task) and optimality (“best description”) (in the best description task). the set-up was exactly the same for both experiment versions. the scenario was that of a shadow play paradigm, telling participants that there is a wizard who likes to play the shadow game with a baby dragon. in this game, various animals go and hide behind the curtain—but some of them may come in front of the curtain later on. the baby dragon has to say who he thinks the shadow belongs to. participants are told that they are supposed to reward the baby dragon with a big apple if what he says is right (right-wrong task) / the best description (best description task) and with a small apple if what he says is wrong (right-wrong task) / not the best description (best description task). importantly, such a contrastive experimental set-up assumes that, in the right-wrong task, subjects will reward both fully informative and underinformative true statements with a big apple, whereas, in the best description task, they will only reward fully informative statements with a big apple, and underinformative true statements will receive a small apple, just like false ones (see table 1): right wrong fully informative underinformative false the best description not the best description table 1: truth, informativity and optimality the experimental materials involve several associated pictures and sentences. each picture has a main silhouette, a small image with the animals in front of the curtain, and a small image with all proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 51 https://doi.org/10.3765/elm https://www.elm-conference.net/ the animals in the game (see figures 1, 2). the small image on the left (all animals) is always present for subjects to easily access the initial situation, without processing difficulties because of memory load (crain & thornton 1998). there are various groups of animals of various colors: a training group of two bunnies and 4 testing groups of three animals each: dog, frogs, cats, cows. we presented participants with 59 sentences in total (3 training sentences, 2x4=8 test sentences, 2x24=48 control sentences containing poate ‘maybe’ or sigur ‘certainly’ presented in a randomized manner, see table 2). the randomization was applied both within the same group of animals and across groups. the test contains a number of sentences balanced between the two epistemic adverbs so as to activate the modal scale and trigger pragmatic readings. the key sentences for implicature detection are highlighted in yellow. all the sentences (except for the practice ones) have the same structure: the adverb poate ‘maybe’ / sigur ‘certainly’ followed by the complementizer cǎ ‘that’ and an embedded sentence. none in front scenario one in front scenario spossible1 underinfo scertain1 optimal spossible2 false scertain2 false spossible3 optimal spossible4 optimal scertain3 overly strong scertain4 overly strong spossible5 false scertain5 false two in front scenario spossible6 underinfo scertain6 optimal spossible7 false scertain7 false table 2: types of utterances tested per scenario 4.2.3. procedure. the experiment started with a training session, followed by the main experiment. in the training session, participants get acquainted with the picture design and practice rewarding on the basis of a bunny shadow picture (see figure 1), where they have to reward the baby dragon with big or small apples. subjects were presented with sentences such as the ones in (1). in the first sentence in (1a), subjects were told which reward to choose (the small apple), while, in the other sentences, they had to choose the reward themselves. the training items were the same in both the right-wrong task and in the best description task. figure 1: item from training session (1) a. este un şoarece /o vacǎ. (false) ‘it is a mouse/a cow.’ b. este un iepuraş. (true/optimal) ‘it is a bunny.’ we will now exemplify the testing session by reference to the group of dogs. scenario 1, the none in front scenario (where all dogs go behind the curtain, see figure 2) ensures that subjects have in mind the set of animals (the referential domain) that is at issue, rather than all the animals in the world, or the animals in the game. sentences (2a) and (2b) are the critical proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 52 https://doi.org/10.3765/elm https://www.elm-conference.net/ conditions. sentence (2c) was a control sentence. subjects are supposed to reason that the animal whose silhouette they see must be a dog, not a cat, not a cow, not a frog, and, hence, choose a small apple as reward. subjects who only consider semantic meaning are expected to reward the dragon with a big apple in conditions (2a) and (2b), and subjects that strengthen the weak epistemic to ‘not certain’ are expected to choose a small apple for (2a) but not for (2b). figure 2: example picture for the none/one/two in front scenarios (2) a. poate cǎ este un cȃine. (underinfo) ‘it is possible that it is a dog.’ b. sigur cǎ este un cȃine (optimal) ‘it is certain that it is a dog.’ c. poate/sigur cǎ este o pisicǎ. (false) ‘it is possible/certain that it is a cat.’ scenario 2, the one in front scenario (where one animal comes back in front of the curtain, in this case, the yellow dog, see figure 2) tests the subjects’ understanding of alternatives, their ability to reason that the situation has two possible outcomes: either the silhouette belongs to the red dog, or it belongs to the blue dog. subjects were expected to choose a big apple for the optimal control statements in (3a) and a small apple for the wrong statements in (3b, c). (3) a. poate cǎ este cȃinele roşu/albastru. (optimal) ‘it is possible that it is the red/blue dog.’ b. sigur cǎ este cȃinele roşu/albastru. (overly strong) ‘it is certain that it is the red/blue dog.’ c. poate/sigur cǎ este cȃinele galben. (false) ‘it is possible/certain that it is the yellow dog.’ scenario 3, the two in front scenario (where two animals are in front of the curtain, see figure 2) tests whether subjects are able to reason that the silhouette can only belong to the blue dog, given that there are two animals in front of the curtain now. sentences (4c, d) were control sentences. subjects who only consider semantic meaning are expected to reward the dragon with a big apple in both (4a) and (4b), and subjects that strengthen the weak epistemic to ‘not certain’ are expected to choose a big apple for (4b) but not for (4a). (4) a. poate cǎ este cȃinele albastru. (underinfo) ‘it is possible that it is the blue dog.’ b. sigur cǎ este cȃinele albastru. (optimal) ‘it is certain that it is the blue dog.’ c. poate cǎ este cȃinele roşu. (false) ‘it is possible that it is the red dog.’ d. sigur cǎ este cȃinele roşu. (false) ‘it is certain that it is the red dog.’ proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 53 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4.3. results. all subjects were included in the analysis of the data given that there were no subjects who made more than 1 error out of 3 in the bunny control items. in order to see the difference in choices between the right-wrong task and the best description task, we performed analyses: (a) on the whole set of data, (b) on subsets of the data: (i) underinformative statements, (ii) (true optimal and false) control statements, and (iii) overly strong statements. 4.3.1. whole data analysis. using r (2018), we computed a logit mixed-effects model with task type and statement type (control, underinformative, overly strong) as fixed effects with treatment coding and item and participant as random effects. the model took the answers to the control condition as the baseline. the results show that the type of task, the overly strong statement type and the underinformative statement type are statistically significant (see table 3). in addition, the interaction between the type of task and the underinformative condition is also significant, but not the interaction between the type of task and the overly strong statement type. table 2: results of a glmer performed on the whole data 4.3.2. subsets of the data. we divide the subset analysis into three parts: scalar implicatures, control statements, and overly strong statements. for scalar implicatures, to determine rates of implicature with more precision, we looked at the corresponding stronger alternative statements with sigur ‘certainly’. subjects were assumed to derive scalar implicatures when they gave a small apple reward to the underinformative statement with poate ‘maybe’ (2a, 4a) and a big apple reward to the stronger alternative statement with sigur ‘certainly’. interestingly, whereas in the right-wrong task, only 29.24% speakers rejected underinformative sentences, with only 14 consistent speakers (i.e., giving more than 5 expected answers out of 8), in the best description task, there were 66.67% scalar answers, with 41 consistent subjects (see figure 3). we ran a logistic regression using a logit mixed-effects model with scalar implicatures as variable, task type and scenario as fixed effects and item and participant as random effects. the results reveal a significant effect for task (β = −3.503, se = 1.664, z = −5.274, p < 0.001) and the interaction between task and scenario (β = 0.393, se = 0.196, z = 2.011, p = 0.044), but no significant effect per scenario (β = −0.203, se = 0.1435, z = −1.417, p = 0.156). figure 3: scalar implicatures per task parameter estimate std. error z p intercept 2.794 0.227 12.293 < 2e-16 *** task type right wrong -0.893 0.193 -4.634 3.58e-06 *** condition overly strong -1.033 0.149 -6.931 4.17e-12 *** condition underinformative -1.423 0.14 -10.096 < 2e-16 *** task type right wrong: condition overly strong 0.097 0.182 0.536 0.592 task type right wrong: condition underinformative -0.913 0.178 -5.121 3.03e-07 proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 54 https://doi.org/10.3765/elm https://www.elm-conference.net/ the results for control sentences (see table 3) reveal more accuracy in the best description task than in the right-wrong task, indicating that the best description encourages participants’ attention. we conducted a logistic regression using a mixed-effects model with task and truth as fixed effects and random by-item and by-participant slopes. the results reveal a significant effect per task (β = −1.164, se = 0.199, z = −5.841, p < 0.001), truth (β = −0.904, se = 0.105, z = −8.56, p < 0.001), as well as the interaction between task and truth (β = 0.265, se = 0.129, z = 2.052, p = 0.004). accuracy best task right-wrong task optimal control sentences false control sentences 84.02% 92.27% 78.83% 80.95% table 3: accuracy in control sentences per task interestingly, in the case of overly strong statements, there was also more accuracy with the best description task (35.67%) than with the right-wrong task (22.17%), see figure 6. a logistic regression using a mixed-effects model with task as a fixed effect and item and participant as random effects reveals a significant task effect (β = −2.877, se = 0.642, z = −4.482, p < 0.001), while including random by-item and by-participant slopes leads to near significance per task (β = 2.607, se= 2.607, z = 1.914, p = 0.055). figure 6: yes to overly strong sentences per task 5. discussion. the whole set analysis indicates that participants behave more accurately with underinformative statements in the best description task. however, we decided to also do subset analyses of the data, given that awarding the dragon with a small apple for underinformative statements is not necessarily an indication of scalar implicatures. to see this, note that there are at least two possible reasons for rejecting the underinformative sentence (4a), either (a) the participant thinks it is impossible for the silhouette to be the blue dog, in which case he/she would also reject the stronger alternative, or (b) the subject thinks it is actually certain, not just possible that it is the blue dog, in which case he/she would accept ‘it is certain that it is the blue dog’. since we did not ask subjects why they gave certain answers-in order to keep the test relatively short, both options are possible in the current experimental set-up. thus, we decided to analyze whether participants derive implicatures by looking at the underinformative sentences (in both scenarios) and at the corresponding stronger alternative statements with certain(ly). we noticed four different patterns of responses in our participants (see table 4): a logical pattern, for subjects who accepted both optimal and underinformative statements, a pragmatic pattern, for subjects who rejected the underinformative statement, but accepted the optimal one, a cautious pattern, for subjects who accepted the underinformative statements, but not the optimal ones with sigur ‘certainly’, and an erroneous pattern, for subjects who rejected both the underinformative and optimal statements. proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 55 https://doi.org/10.3765/elm https://www.elm-conference.net/ table 4: subjects’ patterns of response in underinformative and optimal true sentences interestingly, there is no difference between the proportions of scalar implicatures derived with underinformative sentences belonging to the none in front scenario or the two in front scenario. evaluating whether the silhouette belongs to a certain animal class or is an animal of a certain color leads to similar results (see table 5), suggesting pragmatic consistency across scenarios. table 5: subjects’ patterns of responses per task and scenario the results from our right-wrong task with epistemic adverbs reveal quite low rates of implicatures (29.24%). in contrast, the best description task leads to higher implicature rates (66.67%). on the one hand, it is unclear whether the low rates in the right-wrong task could be explained through the more complex nature of the scalar items tested (epistemic adverbs), as well as language-specific facts related to romanian (the fact that epistemic adverbs select full cps, for instance). we would expect these factors to affect implicature-derivation in the best description task as well, not just in the right-wrong task. on the other hand, both the best description task and the right-wrong task were implemented as a binary task, so the results cannot be understood in terms of an opposition between binary and ternary (as in katsos & bishop 2011). rather, the essential aspect seems to be related to how task instructions model participants’ attention: the right-wrong task encourages adults to pay attention to the truth value of the statements they hear, whereas the best description task encourages adults to pay attention to informativity. as far as the control statements are concerned, the results again suggest better accuracy with the best description task. importantly, the fact that participants had lower accuracy on the true statements suggests there was no yes bias, but rather a tendency to place ‘bets’ on certain animals and reject statements about the possible presence of other animals behind the curtain. in the case of overly strong statements, both tasks had participants who rewarded dragons with big apples (using sigur ‘certainly’ where the weaker poate ‘maybe’ was optimal). the quite high number of big apple rewards for overly strong statements is unexpected given that such statements are false, not optimal. while this could be due to inattention, the higher accuracy in the control statements sheds doubt upon such an explanation. we believe that such answers actually reflect a tendency to place a ‘bet’ on one of the animals when it is yet unknown what animal lies behind the curtain, a tendency which is encouraged by the present task. the best description task taps into subjects’ awareness that they should not place a bet on a certain outcome, hoping that it is true, but rather evaluate whether the statement they hear corresponds to the situation in the best it is possible that it is x it is certain that it is x pattern of response best task right-wrong task big apple small apple big apple small apple big apple big apple small apple small apple logical pragmatic cautious erroneous 24.6% 66.67% 1.38% 7.34% 46% 29.24% 11.7% 13.06% patterns of responses none in front scenario two in front scenario best task right-wrong task best task right-wrong task pragmatic logical cautious erroneous 68.65% 27.38% 1.19% 2.77% 27.34% 53.9% 5.07% 14% 64.68% 21.28% 1.58% 11.9% 31.25% 38.28% 18.36% 12.1% proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 56 https://doi.org/10.3765/elm https://www.elm-conference.net/ way possible or not. thinking whether the baby dragon described the situation best makes it clearer that they are not dealing with a guessing game, but rather a reward game based on evidence. 6. conclusion. the results show a significant task effect in the derivation of scalar implicatures with poate ‘maybe’ when comparing the truth value judgment (the right-wrong task) to the optimality judgment (the best description task). interestingly, accuracy seems to be better in the best description task even with control statements and overly strong statements, which suggests that asking adults to reward characters depending on whether their statement is the best description or not makes adults more attentive. in contrast, there may be more tolerance with respect to what they consider right and wrong. this is very much in line with the katsos & bishop (2011) language acquisition findings. katsos & bishop (2011) argue that the right-wrong binary task masks children’s sensitivity to underinformativeness due to their pragmatic tolerance. in other words, children tend to consider underinformative statements true, but they realize underinformative statements are not optimal. this becomes obvious in their ternary task, where children reward optimal, underinformative, and false statements differently. in the right-wrong task, there were low rates of implicatures, whereas, in the best description task, there were high implicature rates. this indicates that adults are sensitive to task instructions. importantly, testing adults’ interpretation of the stronger alternatives to the underinformative statements allows us to make an informed decision about whether speakers derived scalar implicatures or not. in this way, we can evaluate not only adults’ sensitivity to underinformativeness in the best description task, but their actual ability to derive implicatures. the current experiments thus show the importance of methodology in research: (even linguistically naïve) adults generate implicatures only when asked the adequate question. for these reasons, we recommend using test questions about optimality, especially considering previous research where inferences were not robust. nevertheless, an important point is in order: deciding whether participants derive implicatures implies establishing whether they consider a certain statement true yet underinformative. while the right-wrong task masks sensitivity to underinformativeness (since the statements rewarded with big apples are either optimal or underinformative), the best description task, which was implemented as a binary task as well, masks adults’ truth evaluations (since the statements rewarded with small apples are either underinformative or false). hence, it is extremely important to either ask questions about participants’ reasons for giving a certain answer or include control sentences that evaluate whether participants accepted the stronger alternatives (of underinformative statements)-the latter represents the strategy we adopted in our experiments. another option is to resort to ternary tasks, which test participants’ understanding of optimal, underinformative, and false statements at once, through a three-valued reward system. references begg, ian. 1987. ‘some’. canadian journal of psychology 41:62–73. https://doi.org/10.1037/h0084147. benz, anton & nicole gotzner. 2014. embedded implicatures revisited: issues with the truthvalue judgment paradigm. in j. degen, m. franke, & n. d. goodman (eds.), proceedings of the formal & experimental pragmatics workshop, 1-6. tübingen. benz, anton & nicole gotzner. in press. embedded implicature: what can be left unsaid? linguistics & philosophy. bleotu, adina c. 2019. what colouring can tell us about the acquisition of scalar items in child romanian. osf. september 29. osf.io/bwrvt. proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 57 https://doi.org/10.3765/elm https://www.elm-conference.net/ chemla, emmanuel & benjamin spector. 2011. experimental evidence for embedded scalar implicatures. journal of semantics 28(3): 359–400. https://doi.org/10.1093/jos/ffq023. clifton jr, charles & chad dube. 2010. embedded implicatures observed: a comment on geurts and pouscoulous (2009). semantics and pragmatics 3(7). 1–13. http://dx.doi.org/10.3765/sp.3.7. crain, stephen & cecile mckee. 1985. acquisition of structural restrictions on anaphora. proceedings of north eastern linguistic society (nels) 15. 94-110. crain, stephen & rosalind thornton. 1998. investigations in universal grammar: a guide to experiments on the acquisition of syntax and semantics. cambridge, ma: mit press degen, judith & michael k. tanenhaus. 2011. making inferences: the case of scalar implicature processing. in l. carlson, c. hӧlscher & t. shipley (eds.), in proceedings of the 33rd annual conference of the cognitive science society. 3299–3304. foppolo, francesca, maria teresa guasti, & gennaro chierchia. 2012. scalar implicatures in child language: give children a chance. language learning and development 8. 365-394. https://doi.org/10.1080/15475441.2011.626386. geurts, bart & nausicaa pouscoulous. 2009. embedded implicatures?!? semantics and pragmatics 2(4). 1–34. http://dx.doi.org/10.3765/sp.2.4. gotzner, nicole & anton benz. 2018. the best response paradigm: a new approach to test implicatures of complex sentences. frontiers in communication 2(21). 1-13. https://doi.org/10.3389/fcomm.2017.00021. heizmann, tanja. 2006. acquisition of deontic and epistemic readings of must and müssen. in tanja heizmann (ed.), university of massachusetts occasional papers in linguistics (umop) 34: current issues in language acquisition. amherst, ma: glsa, umass amherst. hirst, william & joyce weil. 1982. acquisition of epistemic and deontic meaning of modals. journal of child language, 9(3). 659–666. https://doi.org/10.1017/s0305000900004967. katsos, napoleon & dorothy bishop.2011. pragmatic tolerance: implications for the acquisition of informativeness and implicature. cognition 120 (1). 67-81. https://doi.org/10.1016/j.cognition.2011.02.015. noveck, ira. 2001.when children are more logical than adults. cognition 78(2). 165-188. https://doi.org/10.1016/s0010-0277(00)00114-1. noveck, ira a., simin ho & maria sera. 1996. children's understanding of epistemic modals. journal of child language 23 (3): 621-643. https://doi.org/10.1017/s0305000900008977. ozturk, ozge & anna papafragou. 2015. the acquisition of epistemic modality: from semantic meaning to pragmatic interpretation. language learning and development 11 (3). 191-214. https://doi.org/10.1080/15475441.2014.905169. papafragou, anna & julien musolino. 2003. scalar implicatures: experiments at the semantics pragmatics interface, cognition 86(3). 253-282. https://doi.org/10.1016/s00100277(02)00179-8. rosch, eleanor. 1973. natural categories. cognitive psychology 4(3). 328–50. https://doi.org/10.1016/0010-0285(73)90017-0. rosch, eleanor. 1975. cognitive representations of semantic categories. journal of experimental psychology 104(3):192–233. https://doi.org/10.1037/0096-3445.104.3.192. van tiel, bob. 2013. embedded scalars and typicality. journal of semantics 31(2).147-177. https://doi.org/10.1093/jos/fft002. zehr, jeremy & florian schwarz. 2018. penncontroller for internet based experiments (ibex). https://doi.org/10.17605/osf.io/md832. proceedings of elm 1: 047-058, 2021 adina camelia bleotu, anton benz and nicole gotzner: where truth and optimality part. experiments on implicatures with epistemic adverbs. 58 https://doi.org/10.3765/elm https://www.elm-conference.net/ or not alternative questions, focus and discourse structure maribel romero, erlinde meertens & andrea beltrama* abstract. or not alternative questions like are you coming or not? give rise to socalled ‘cornering effects’ (biezma 2009), consisting of two parts: (i) they cannot appear discourse-initially, and (ii) they do not allow for follow-up questions. building on recent experimental data (beltrama, meertens & romero 2020), the present paper raises problems for current analyses (biezma 2009, biezma & rawlins 2012, 2018), reframes the second part of cornering as not specific to naqs but as a general constraint on questions in general, and develops a novel proposal for the first part of cornering. the key ingredients of the new proposal are the intrinsic focus structure of or not questions and its effects on discourse trees. keywords. alternative question; polar question; cornering; focus; polarity; negation; discourse structure 1. introduction. consider the polar question (pq) in (1) and the or not alternative question (negative alternative question, naq) in (2). the two interrogative forms raise the same issue. in terms of possible answers in the hamblin-style denotation, they denote the set (3) containing the prejacent proposition and its negation. in terms of groenendijk & stokhof’s (1984) pragmatic partitions, they trigger the same partition (4). that is, the two forms are informationally equivalent: (1) are you giving a talk at elm? pq (2) are you giving a talk at elm or not? or-not-altq /negative altq (naq) (3) { lw. you are giving a talk at elm in w, lw.¬(you are giving a talk at elm in w) } (4) nevertheless, pqs and naqs differ in their use-conditions. a well-known difference concerns so-called cornering effects, first discussed by biezma (2009).1 intuitively, naqs convey a sense of insistence, as if the speaker was trying to “corner” the addressee into answering the question. pqs, in contrast, do not engender this effect. more concretely, biezma (2009) formally characterizes cornering effects as two constraints that apply to the discourse distribution of naqs but not of pqs. the first constraint –part 1 of cornering in (5a)– precludes naqs but not pqs from discourse-initial position, as illustrated in (6). the second constraint –part 2 of cornering in (5b)– prohibits follow-up questions to naqs but not to pqs, as exemplified in (7): (5) cornering effects: i. part 1: pqs can occur discourse initially whereas naqs cannot: (6) ii. part 2: pqs allow for follow-up questions whereas naqs do not: (7) * we thank the audience of experiments in linguistics meaning (elm 1) for questions and comments and daniel goodhue for further discussion. this research has been supported by the dfg project ro 4247/4-2. authors: maribel romero (university of konstanz, maribel.romero@uni-konstanz.de), erlinde meertens (university of konstanz, erlinde.meertens@uni-konstanz.de) and andrea beltrama (university of pennsylvania, andrea.beltrama@gmail.com). 1 for differences between pqs and naqs in terms of illocutionary acts and alike, see bolinger (1978). lw. you are giving a talk at elm in w lw.¬(you are giving a talk at elm in w) proceedings of elm 1: 249-260, 2021 c©2021 maribel romero, erlinde meertens and andrea beltrama published by the lsa with permission of the author(s) under a cc by license. 249 https://doi.org/10.3765/elm https://www.elm-conference.net/ (6) a. s: hi dad! i’m hungry. ! are you making pasta? b. s: hi dad! i’m hungry. # are you making pasta or not? (# unless the issue was discussed before) (7) s: are you making pasta? (based on biezma 2009) a: (silence and dubitative faces) s: are you making pasta? a: hmm… s: ! are you making pasta or not? a: (silence and dubitative faces) s: # are you making pasta? biezma’s (2009) intuitive idea is that naqs are only adequate in a particular position in the discourse tree (d-tree) representing the question under discusion (qud) structure of discourse (roberts 1996/2012): naqs close up an entire line of inquiry or discourse strategy in a d-tree. a sample d-tree is given in (8). given this mandatory position in the d-tree, naqs are (i) infelicitous discourse initially (cornering part 1), since there is no line of inquiry to close off, and (ii) do not allow for follow-up questions or sub-questions (cornering part 2), since by definition they close the strategy they are part of: (8) what are you making for dinner? are you making pasta? are you making fish? are you making a stew? … are you making pasta or not? … the question is, then, why naqs must mandatorily occupy this position in the d-tree. the literature currently offers two (subsequent) answers to this question. under approach a (biezma 2009, biezma & rawlins 2012), this position in the d-tree results from the logical exhaustivity of the disjuncts in naqs. under approach b (biezma & rawlins 2017), it follows from the combination of logical exhaustivity and bundling of alternatives under a negative description.2 these two approaches aimed at deriving the contrast in discourse distribution between pqs and naqs despite their informativity equivalence. interestingly, besides pqs and naqs, there is a third question form that gives rise to the very same issue or partition: complement alternative questions (caqs), in which the main predicate in the second disjunct is a (lexical or phrasal) complement of the predicate in the first disjunct. to see this, consider the triple in (9). the three question forms give rise to same set of possible answers in the hamblin-style denotation (10) and to the same pragmatic partition (11). hence, they can be seen as informationally equivalent: (9) a. pq: is the light on? b. naq: is the light or not? c. caq: is the light on or off? (10) { lw. the light is on in w, lw.¬(the light is on in w) (= lw. the light is off in w) } 2 biezma & rawlins (2017) are mostly concerned about bundling under what in questions like (i). they only tackle naqs in passing. (i) are you cooking pasta or what? cornering part 1 cornering part 2 proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 250 https://doi.org/10.3765/elm https://www.elm-conference.net/ (11) as we will discuss, beltrama, meertens & romero (2020) experimentally tested pqs, naqs and caqs and found that each question type displays a different pattern in terms of cornering effects, raising issues for both existing approaches a and b. in the present paper we take a new approach the two constraints on the distribution of naqs underlying cornering. the idea, in a nutshell, is the following. first, we follow beltrama et al. (2020) in suggesting that part 2 of cornering is not an inherent property of naqs, but rather reflects a broader constraint on discourse moves, from which naqs can indeed be exempt once the proper discourse conditions are met. second, we develop a novel account of part 1 of cornering which partially retains biezma’s appealing idea that naqs but not pqs –and not caqs either– are limited to a certain position in the discourse tree. we derive this limitation from two ingredients: polarity focus and discourse structure. the paper is structured as follows. section 2 reviews previous approaches, summarizes the main results from beltrama et al. (2020) and raises additional challenges for existing approaches. section 3 presents the key ingredients of the proposal on part 1 of cornering. section 4 applies the proposal to pqs, naqs and caqs. section 5 concludes. 2. comparing pqs, naqs and caqs. 2.1. previous approaches and their predictions for caqs. recall biezma’s (2009) idea: the limited distribution of naqs –parts 1 and 2 of cornering in (5)– stems from the fact that, given some property of their form, they must occupy a particular position in the d-tree, as in (8). according to approach a (biezma 2009, biezma & rawlins 2012), the determining property of naqs is the fact that their disjuncts exhaust together the logical space of possibilities. more specifically, the contrast between pqs and naqs is derived as follows. pqs asks about one alternative (e.g. are you making pasta?) while leaving other potential implicit alternatives open (e.g., are you making fish?, are you making a stew, etc.), as in the d-tree (8). with this, the speaker is granting the addressee a high degree of freedom in discourse. in contrast, naqs, by presenting two alternatives that exhaustify the possibility space, force the listener to choose one of them, crucially restricting their room for maneuvering. it is argued that this blatant limitation of the hearer’s maneuvering space is too forceful to begin a strategy, thus prohibiting naqs in discourse initial position (cornering part 1); and that this forcefulness signals that the strategy comes to an end, hence disallowing follow-up questions (cornering part 2). what predictions does approach a make for caqs? like naqs and unlike pqs, the disjuncts in caqs exhaust together the logical space of possibilities. thus, naqs are predicted to pattern like naqs and unlike pqs for both parts of cornering, as noted in (12): (12) predictions for caqs according to approach a: a. # discourse initial (cornering part 1) b. # follow-up questions (cornering part 2) according to approach b (biezma & rawlins 2017), naqs occupy that particular position in the d-tree because their disjuncts exhaust together the logical space of possibilities and, crucially, other than the first alternative p, all other alternatives q, r, s … are bundled in the second disjunct under a negative description, namely, under the description of not being the first alternative. this lw. the light is on in w lw.¬(the light is on in w) (= lw. the light is off in w) proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 251 https://doi.org/10.3765/elm https://www.elm-conference.net/ is argued to derive the discourse behavior of pqs and caqs as follows. on the one hand, pqs involve no exhaustivity of the logical space and no bundling of alternatives p, q, r, s, etc. thus, no discourse constraints arise. on the other hand, bundling in general and bundling under negation in particular has several consequences for naqs. first, naqs convey a relevance asymmetry due to bundling under negation: the speaker signals that they are merely interested in the content proposition p, thus eliminating the remaining alternatives q, r, s… from future relevance to the discourse. this imbues caqs with a peremptory feeling, which is appropriate for ending a strategy (cornering part 2). second, naqs provide no cues as to what alternatives are bundled under negation. but in discourse initial contexts it typically matters what the other alternatives are. this undermines naqs’ felicity in discourse-initial position (factor (a) for cornering part 1). third and finally, naqs involve bundling several alternatives together. bundling has a cost (it is an accommodation move) and thus needs a motivation. no motivation is available in discourse initial contexts, jeopardizing again their felicity in this position (factor (b) for cornering part 1). what predictions follow from approach b for caqs? like pqs and unlike naqs, caqs involve no bundling of alternatives. for instance, in example (9c), there are only two states that the light might be in: on and off. this means that the second disjunct the light is off is not a cover term bundling together several possible states. therefore, naqs are predicted to pattern like pqs and unlike naqs for both parts of cornering, as in (13): (13) predictions for caqs according to approach b: a. ! discourse initial (cornering part 1) b. ! follow-up questions (cornering part 2) the predictions of the two existing approaches for caqs are summarized in table 1. approach a approach b pq naq caq pq naq caq part 1: discourse initial ! # # ! # ! part 2: with follow-up questions ! # # ! # ! table 1: predictions of existing approaches a and b 2.2. experimental findings from beltrama et al. (2020). in two rating experiments, beltrama et al. (2020) compare the behavior of pqs, caqs and naqs with respect to part 1 and part 2 of cornering respectively. to test part 1 of cornering, they compare the naturalness of these three questioning strategies in discourse initial position, asking participants to provide a judgment on a 7 point scale (1 = completely unnatural; 7 = perfectly natural; see beltrama et al. 2020, section 4 for details on the materials, procedure and analysis). two aspects of their findings are especially relevant to our purposes (see beltrama et al. 2020: for a more exhaustive discussion). first, naqs turned out to be significantly less natural than pqs, replicating biezma’s (2009) intuitive contrast between pqs and naqs discourse-initially (see (6) above). second, contrary to naqs, caqs were as felicitous as pqs discourse-initially, and significantly more felicitous than naqs. taken together, these results fail to support the prediction (12a) of approach a with respect to part 1 of cornering, while providing support to the prediction (13a) of approach b. to compare the behavior of naqs and caqs with respect to part 2 of cornering, and therefore shed light on each of these question's ability to license follow-up moves, they carried out a further rating study, which came in two different versions. in this study, they constructed proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 252 https://doi.org/10.3765/elm https://www.elm-conference.net/ dialogues containing a sequence of three questions. the first question was always a pq; the second question was either a naq or a caq; and the third question was either a question identical to the opening pq, or different from it –i.e., a pq question asked with emphatic tone in one version of the experiment or a wh-question in another version of the experiment. a sample item is provided in (14) (see beltrama et al. 2020, section 5 for details about materials, procedure and analysis): (14) sophia and rachel are about to play chess. the following exchange ensues. sophia: are you playing with black? pq rachel: well, i can't wait to play sophia: are you playing with black { or not? / or white? } naq/caq rachel: i'm going to crush you! sophia: are you playing with black? / are you playing with black? sophia: are you playing with black? / what color do you want to play with? the findings from this study highlight three crucial takeaways. first, contra biezma’s (2009) original characterization and contra approaches a and b, naqs do not uniformly disallow followup questions. in particular, when the intended follow-up question (3rdq in the dialog) is identical to the original pq question (1stq in the dialog), the ratings are significantly lower than when the intended follow-up question (3rdq) is different form the original pq (1stq) –i.e., it is an emphatic pq in study 2a or a wh-question in study 2b. second, caqs pattern exactly like naq when it comes to the ability to license follow-up questions in discourse: similar to what is observed for naqs, follow-up questions to caqs are considerably more natural when they have not been used in discourse yet than when they have already been used. third, no effect of the previous question type was found on the naturalness of follow-up questions: follow-up questions to naqs and caqs were equally natural when they had not been deployed in previous discourse yet; and equally degraded when they had been used already. this contradicts the prediction (13b) of approach b, which expects difference tolerance for follow-ups between naqs –due to bundling– and caqs – since they involve no bundling. 2.3. empirical support for cornering: taking stock. table 2 summarizes the results of the experiments and compares them with the predictions made by the two approaches to cornering. approach a approach b exp results pq naq caq pq naq caq pq naq caq part 1: discourse initial ! # # ! # ! ! # ! part 2: with follow-up questions ! # # ! # ! ! # / ! # / ! table 2: predictions of existing approaches a and b plus results of exp1 and exp1a/b the findings from these experiments support two important conclusions. first, while the degraded status of naqs in discourse-initial position confirms that they are indeed subject to part 1 of cornering, the same degrading is not observed for caqs, suggesting that this particular question type is not subject to this part of cornering. second, the results from the studies exploring the behavior of pqs, naqs and caqs with respect to part 2 of cornering show that the key factor determining the un/acceptability of follow-up questions is not the question form itself (2ndq) preceding the intended follow-up, but rather the relation between the original question (1stq) and the final follow-up question (3rdq): using the same form leads to infelicity whereas using a different form leads to felicity. this leads beltrama et al. (2020) to conclude that part 2 of cornering is not a distinctive propery of naqs per se but rather the result of a more general proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 253 https://doi.org/10.3765/elm https://www.elm-conference.net/ constraint on discourse. they suggest a constraint along the lines of (15), applicable not just to naqs but to questions (or strategies) in general: (15) * repeat: do not use a question form that has been unsuccessfully used and abandoned in previous discourse. for the present paper, this means that cornering part 2 does not need to be derived from any property arising from the specific form of naqs, but can instead be captured in terms of broader constraints governing when a discourse move is pragmatically admissible. what remains to be derived, however, is part 1 of cornering, which we now turn to discuss. 2.4. against approach b for cornering part 1. should we maintain approach b to derive the infelicity of naqs in discourse initial position? we present two arguments against this move, one for each of the factors (a) and (b) impacting discourse initiality according to approach b. we start with factor (a), bundling under a negative description. this factor has been argued to lead to part 1 of cornering in that it obscures what the bundled alternatives q, r, s… are and such alternatives typically matter in discourse initial contexts. we note, though, that there exist discourse initial contexts which do allow for question forms that deliberately signal that the only relevant live possibility is the content proposition p and that no other alternative q, r, s… matters. this is the case e.g. of pqs with a falling final contour like (16) (h*l-l% in tobi notation) (bartels 1999, westera 2017; see also roelofsen & van gool 2010, biezma & rawlins 2012). in such permissive discourse initial contexts, the effects of bundling under negation should bring no infelicity. however, naqs are still infelicitous in such contexts, witness (17): (16) u.s. immigration officer to the next traveler in line: are you a u.s. citizenh*l-l%? (17) u.s. immigration officer to the next traveler in line: # are you a u.s. citizen or not? we turn to factor (b), the cost of bundling alternatives. bundling is a costly move, which means that a motivation is required to use a bundled question strategy like (19) instead of its unbundled counterpart (18). it is argued that, since discourse initial contexts do not (typically?) provide such motivation, question forms that involve bundling are infelicitous in that position. (18) … are you making pasta? are you making fish? are you making broccoli? {lw. you make pasta in w} {lw. you make fish in w} {lw. you make brocc. in w} (19) … are you making pasta or not? {lw. you make pasta in w, lw. you make fish in w ú you make broccoli in w} but consider a minimal pair of a caq and the corresponding naq, e.g. (9c) and (9b). in order to (correctly) predict that caqs are felicitous discourse initially, approach b would need to assume that using the caq strategy (21) instead of the pq strategy (20) requires no motivation. but, then, naqs with only one alternative other than p are predicted to behave exactly like their caqs counterparts in requiring no motivation, since the caq-version and the naq-version denote the proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 254 https://doi.org/10.3765/elm https://www.elm-conference.net/ very same set of propositions (since [lw. the light is off in w] = [lw. ¬(the light is on in w)]) and thus introduce the very same strategy (21)/(22). however, our experimental results showed a contrast here: the caq-versions are felicitous discourse initially while the naq-versions are not. (20) … is the light on? is the light off? {lw. the light is on in w} {lw. the light is off in w} (21) … (22) … is the light on or off? is the light on or not? {lw. light is on in w, lw. light is off in w} {lw. light is on in w, lw.¬(light is on in w)} 3. ingredients of the proposal. our proposal for part 1 of cornering consists of two main ingredients. first, naqs carry f(ocus)-marking on the polarity heads of the two disjuncts. second, f-marking must be suitably licensed by the previous discourse, as in e.g. roberts’ (1996/2012) and büring’s (2003) discourse structure framework. we introduce each ingredient in turn. 3.1. naqs carry focus-marking on the polarity head. it is known that q/a pairs like (23) require the f-marked element in the answer to match the wh-element in the question, and that contrastive structures like (24) require the f-marked –and the c(ontrastive) t(opic) marked– constituents in the two conjuncts to be parallel (rooth 1992, büring 2003): (23) q: who did betty see yesterday? a: betty saw alif yesterday. a’: # betty saw ali yesterdayf. (24) a. bettyct saw alif yesterday and terryct saw marthaf yesterday. b. # bettyct saw alif yesterday and terryct saw ali/him todayf. using q/a and contrastive structures as diagnostics, we can see that focal stress on the lexical verb –or on the main predicate– is ambiguous in (at least) two ways: it may indicate f-marking on the stem, as disambiguated by the q/a (25) and the constrastive structure (26), or it may signal fmarking on the polarity head, as disambiguated by the q/a (27) and the constrast structure (28):3 3 focus stress on an inserted auxiliary might be considered to signal f-marking on the polarity as well: (i.b’) and (ii). two notes about this. first, some authors consider this construction to be degraded as an answer to a pq, as in (i.b’), unless the pq is biased towards the negative answer, as in (ii) (gutzmann et al. 2020, wilder 2013); other authors find (i.b’) perfectly acceptable (goodhue 2020). we side with the former for the spanish translation in (iii). second, some authors take the construction to signal the presence of a verum operator (gutzmann et al. 2020); others argue for f-marking on the polarity (goodhue 2020). here we tentatively side with the latter for spanish, since the locution de verdad ‘of truth’ signaling verum and the spanish construction at issue do not give rise to parallel intuitions in (iv). given this additional complexity, we leave this construction aside. (i) a: did chris submit her paper yesterday? (ii) a: did chris really submit her paper yesterday? b: yes, she submitted her paper. b: yes, she did submit her paper. b’: % yes, she did submit her paper. (iii) sí que entregó su artículo. (iv) a. de verdad entregó cristina su artículo? yes that submitted her article. of truth submitted c her article ‘she did submit her article.’ b. sí que entregó cristina su artículo? yes that sumbitted c her article proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 255 https://doi.org/10.3765/elm https://www.elm-conference.net/ (25) q: what did james do with his parents last christmas? a: he visited them. lf of (25a): [he pol+ visitf them] (26) katect loves monty python sketches and paulct despises them. lf: [katect pol+ lovesf mp sketches and paulct pol+ despisesf them] (27) q: did james visit his parents last christmas? a: (yes,) he visited them. lf of (27a): [he pol+f visit them] (28) katect loves monty python sketches and paulct doesn’t. lf: [katect pol+f loves mp sketches and paulct pol-f love them] when it comes to altqs in general, it has been noted that the two disjuncts must bear focus stress on parallel constituents (bartels 1999, han & romero 2004, truckenbrodt 2013), as in (29): (29) a. did you see alif yesterday or rashmif? b. # did you see ali yesterdayf or rashmif? just like in the cases above, focal stress of the lexical verb –or main predicate– in altqs is ambiguous between f-marking on the stem or on the polarity. this ambiguity is resolved in altqs via the second disjunct, yielding f-marking on the stem in (30) and on the polarity in (31): (30) does kate love monty python sketches or does she despise them? lf: [ q [ [ kate pol+ lovef mt sketches] or [she pol+ despisesf them] ] ] (31) does kate love monty python sketches or not? lf: [ q [ [ kate pol+f love mt sketches] or [she pol-f love them] ] ] using, thus, the second disjunct as diagnostic of the underlying location of f-marking, we arrive at the following conclusion. while the sentences in (32) all raise the same issue, they differ in f-marking possibilities: the pq (32a) is ambiguous between f-marking on the verbal stem or on the polarity; the naq (32b) unambiguously involves f-marking on the polarity; and the caq (32c) unambiguously features f-marking on the predicate’s stem. we argue that it is this difference in f-marking that determines their felicity in discourse initial position. (32) a. pq: is the light on? [q pol+ [the light be-onf]] or [q pol+f [the light be-on]] b. naq: is the light on or not? [q pol+f [the light be-on]] c. caq: is the light on or off? [q pol+ [the light be-onf]] 3.2. roberts (1996/2012) and büring (2003) discourse structure framework. following roberts (1996/2012), the structure infostrd of a discourse d includes a hierarchically ordered set of implicit or explicit moves (questions and answers, viewed as semantic objects), as in (33): (33) 1. 'who{john,paul} call whom{amy,betty}?' a. 'who called amy?' i. 'did john call amy?' ii. 'did paul call amy?' b. 'who called betty?' i. 'did john call betty?' proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 256 https://doi.org/10.3765/elm https://www.elm-conference.net/ ii. 'did paul call betty?' building on roberts (1996/2012) and schwarzschild (1999), büring (2003) proposes that an utterance u can serve as the realization of a given move m in a d-tree d if and only if u is congruent with the d-tree hosting m in the way described by the conditions (34)-(36): (34) givenness condition: [slightly simplified here] an utterance [cp c0 [ip b]] can realize a move m within a d-tree d only if, for every constituent c in b, there is a salient antecedent a in the explicit realization of a move a preceding m in d such that: a. if c is of type e, then c and a corefer; b. otherwise, modulo $-type shifting, a entails the existential f-closure of c. (35) existential f-closure of a constituent c =df the result of replacing f-marked phrases in c with variables and $-binding them. (36) $-type shifting of a constituent c =df the result of existentially binding unfilled arguments of c. within this framework, one can account not just for q/a pairs, but also for q…q sequences like (37), as follows. we start with the explicitly realized preceding move –the question who called amy? in (37)– and its $-type shifting (38). now we want to realize the question move ‘did john call amy?’. if we realize this move using the corresponding interrogative with f-marking on john, each constituent c in this interrogative will be entailed (modulo $-type shifting and existential fclosure) by the $-type shifting in (38), as sketched in (39). in contrast, if we realize this move using f-marking on amy, the constituent(s) (40c,d) in this interrogative will fail to be entailed (modulo $-type shifting and existential f-closure) by (38):4 (37) a. who called amy? did johnf call amy? / did johnf or paulf call amy? b. #who called amy? did john call amyf ? / did john call amyf or suef? (38) $-type shifting of who called amy?: lw. $x[callw(x,amy)] (39) [ip [johnf call amy]] (40) [ip [john call amyf]] a. [np amy]: coreferential a. [np amyf]: lw.$p$y[p(y)(w)] b. [vp call amy]: lw.$x[callw(x,amy)] b. [vp call amyf]: lw.$x$y[callw(x,y)] c. [np johnf ]: lw.$p$x[p(x)(w)] o c. [np john]: not coreferential d. [ip johnf call amy]: o d. [ip john call amyf]: lw.$x[callw(x,amy)] lw.$y[callw(john,y)] in the following section, we combine the two ingredients introduced in this section and apply them to pqs, caqs and naqs to derive their behavor with respect to cornering part 1. 4 one could in principle think that the foci in the disjuncts of did john or paul call amy? in (37a) and of did john call amy or sue? in (37b) can license each other, having the elliptical structures in (i) and using rooth (1992) without appeal to qud structure (as in e.g. han & romero 2004). this, however, would not derive the contrast between the felicitous q…q sequence (37a) and the infelicitous (37b). (i) a. did [john call amy] or [paul call amy]? b. did [john call amy] or [john call sue]? proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 257 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4. applying the proposal to pqs, naqs and caqs discourse initially. recall the crucial empirical result in table 2: in discourse initial contexts, pqs and caqs are felicitous while naqs are not. a sample triple is given in (41) below, where focal stress is now explicitly marked: (41) t(ess): (guess what!) jane got a new car! b(ob): a. ! is it automatic? pq b. # is it automatic or not? naq c. ! is it automatic or a stick shift? caq a partial discourse structure for (41) is given in (42). the utterance (41t) has realized move (1.1.a), underlined in the d-tree (42), and has provided the proposition (43), against which the entailment patterns of the upcoming move will be checked. now speaker b wants to realize move (1.2.a.i), boldfaced in the d-tree (42). the options for the overt realization of this move are the interrogative forms in (41a,b,c). we will see each in turn. (42) 1. 'what is new with jane?' 1.1. 'did jane get a new car?' a. 'jane got a new car.' b. 'jane did not get a new car.' 1.2. 'what properties does jane's new car have?' a. 'what transmission system does jane's new car have?' i. 'is jane's new car automatic?' ii. 'is jane's new car a stick shift?' b. 'what color is jane's new car?' . . . (43) propositional content of a’s utterance jane got a new car! lw. $x [new-carw(x) ù havew(jane,x)] we start with naqs. to the infelicitous (44b) corresponds an lf with f-marking on the polarity heads pol+ and pol-. we see that constituent (45a) is correferential with a previously realized np and that the (tautological) $-type shifted existential f-closure of constituent (45d) is (trivially) entailed by the previous proposition (43). hence, these two constituents satisfy the givenness condition (34). but now consider constituent (45b). its $-type shifted existential fclosure lw.$y[automaticw(y)] is not entailed by the only previously available proposition (43). the same fate awaits constituent (45c), whose $-type shifted f-closure lw.automatic(g(1))(w) is not entailed by the previous discourse. hence, these two constituents violate the givenness condition (34) and, as a result, the naq realization of move (1.2.a.i) in this d-tree is infelicitous: (44) t: jane got a new car! b: # is it automatic or not? lf of (44b): [q [ [pol+f it be-automatic] or [pol-f it be automatic] ] ] (45) [pol+f it be-automatic] / [pol-f it be-automatic] a. [np it1]: coreferential o b. [vp be-automatic]: lw. $y[automaticw(y)] o c. [ip it1 be automatic]: lw. automatic(g(1))(w) d. [ip pol+f [it1 be automatic]] / [ip pol-f [it1 be automatic]]: lw. $z[z(lw'.[automaticw'(g(1))])] proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 258 https://doi.org/10.3765/elm https://www.elm-conference.net/ next, we tackle caqs. they give rise to the felicitous (46b), whose lf has f-marking on the main predicate. the $-type shifted existential f-closure of each constituent is provided in (47). crucially, each of these constituents is either correferential with a previous np or entailed by the previous proposition (43). this renders the caq realization of move (1.2.a.i) in the d-tree felicitous: (46) t: jane got a new car! b: is it automatic or a stick shift? lf of (46b): [q [ [pol+ it be-automaticf] or [pol+ it be-a-stick-shiftf] ] ] (47) [pol+ it be-automaticf] / [pol+ it be-a-stick-shiftf] a. [np it1]: coreferential b. [vp be-automaticf] / [vp be-a-stick-shiftf]: lw. $p$y[p(y)(w)] c. [ip it1 be-automaticf] / [ip it1 be-a-st-shiftf]: lw. $p[p(g(1))(w)] d. [ip pol+ [it1 be-automaticf]] / [ip pol+ [it1 be-a-st-shiftf]]: lw. $p[p(g(1))(w)] finally, we come to the pq realization in (48b), where the prosodic stress naturally falls on automatic. recall from section 3.1 that, when the stress falls on the main predicate, the sentence is ambiguous between f-marking on the polarity head, as in the lf1 below, and f-marking on the predicate stem, as in the lf2. the former structure contains the same problematic constituents (45b,c) that we saw in naqs and, thus, will be ruled out in this discourse initial context. but the parse (48d) gives us the same underlying structure as the first disjunct of the caq version, which we saw in (47) satisfies the givenness condition for all its constituents. the availability of this second parse makes pqs well-suited for discourse initial uses. (48) t: jane got a new car! b: is it automatic? lf1 of (48b): [q [pol+f it be-automatic] ] lf2 of (48b): [q [pol+ it be-automaticf] ] in sum, the structure with f-marking on the main predicate underlying caqs and (a parse of) pqs satisfies givenness in discourse initial contexts like (44), while f-marking on the polarity in naqs does not. this derives the contrast among the different interrogative forms for part 1 of cornering: pqs and caqs are exempt from cornering part 1, while caqs are subject to it. 5. conclusions. the experimental results in beltrama, meertens & romero (2020) led to the following empirical conclusions. first, the inability to appear discourse initially –cornering part 1– characterizes naqs and not pqs and caqs. second, the ban on follow-up questions – cornering part 2– is not an inherent characteristic of naqs; rather, naqs (and caqs) disallow already used interrogative forms as follow-ups but allow novel interrogatives as follow-ups. this posits problems for previous accounts. on the one hand, approach a on exhaustive disjuncts is problematic in view of the experimental results on cornering part 1 and part 2. on the other, approach b on bundling is problematic in view of the experimental results on cornering part 2. additionally, approach b faces conceptual challenges with respect to cornering part 1. following beltrama et al. (2020), we have reframed cornering part 2 as non intrinsic to naqs per se, but as a general constraint on question strategies in general. to capture cornering part 1, we have developed a proposal that maintains biezma’s (2009) intuitive d-tree idea but derives it in a different way. it features two main ingredients: (i) naqs mandatorily carry f(ocus)-marking on the polarity heads of the two disjuncts whereas caqs and proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 259 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a parse of) pqs do not; and (ii) f-marking must be suitably licensed by the previous discourse, e.g. as in roberts (1996/2012) and büring (2003) discourse structure framework. this derives the empirical pattern attested in our experimental results: pqs and caqs can be used in discourse initial contexts while naqs cannot. references bartels, christine. 1999. the intonation of english statements and questions. new york: garland. beltrama, andrea, erlinde meertens & maribel romero. 2020. alternative questions: distinguishing between negated and complementary disjuncts. semantics and pragmatics 13. 1-27. http://dx.doi.org/10.3765/sp.13.5. biezma, maría. 2009. alternative vs polar questions: the cornering effect. proceedings of semantics and linguistic theory (salt) 19. 37–54. https://doi.org/10.3765/salt.v19i0.2519. biezma, maría & k. rawlins. 2012. responding to alternative and polar questions. linguistics & philosophy 35(5). 361-406. https://doi.org/10.1007/s10988-012-9123-z. biezma, maría & kyle rawlins. 2017. or what?. semantics and pragmatics 10(16). 1-44. http://dx.doi.org/10.3765/sp.10.16. bolinger, dwight. 1978. yes-no questions are not alternative questions. in henry hiz (ed.), questions, 87-105. dordrecht: reidel. büring, daniel. 2003. on d-trees, beans and b-accents. linguistics & philosophy 26. 511–545. https://doi.org/10.1023/a:1025887707652. goodhue, daniel. 2020. there is a high-end convertible: polarity focus, contrastive focus, answer focus, givenness. ms. university of maryland. groenendijk, jeroen & martin stokhof. 1984. studies on the semantics of questions and the pragmatics of answers. amsterdam: university of amsterdam dissertation. gutzmann, daniel, katharina hartmann & lisa matthewson. 2020. verum focus is verum, not focus: cross-linguistic evidence. glossa 5(1). 1-48. http://doi.org/10.5334/gjgl.347. hamblin, charles l. 1973. questions in montague grammar. foundations of language 10(1). 4153. https://doi.org/10.1016/b978-0-12-545850-4.50014-5. han, chung-hye & maribel romero. 2004. disjunction, focus and scope. linguistic inquiry 35. 179-217. https://doi.org/10.1162/002438904323019048. roberts, craige (1996/2012). information structure in discourse: towards an integrated formal theory of pragmatics. semantics and pragmatics 5(6). 1-69. http://dx.doi.org/10.3765/sp.5.6. roelofsen, floris & sam van gool. 2010. disjunctive questions, intonation, and highliting. in maria aloni, harald bastiaanse, tikitu de jager & katrin schulz (eds.), logic, language and meaning: selected papers from the 17th amsterdam colloquium, 384–394. heidelberg: springer. rooth, matts. 1992. a theory of focus interpretation. natural language semantics 1. 75–116. https://doi.org/10.1007/bf02342617. schwarzschild, roger. 1999. givenness, avoidf and other constraints on the placement of accent, natural language semantics 7(2). 141–177. https://doi.org/10.1023/a:1008370902407. truckenbrodt, hubert. 2013. an analysis of prosodic f-effects in interrogatives: prosody, syntax and semantics. lingua 124. 131-175. https://doi.org/10.1016/j.lingua.2012.06.003. westera, matthijs. 2017. exhaustivity and intonation: a unified theory. amsterdam: university of amsterdam dissertation. wilder, chris. 2013. english ‘emphatic do’. lingua 128. 142-171. https://doi.org/10.1016/ j.lingua.2012.10.005. proceedings of elm 1: 249-260, 2021 maribel romero, erlinde meertens and andrea beltrama: or not alternative questions, focus and discourse structure. 260 https://doi.org/10.3765/elm https://www.elm-conference.net/ shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic! adina camelia bleotu, anton benz & nicole gotzner* abstract. the current paper employs a novel shadow play paradigm to investigate the semantic knowledge and pragmatic ability of romanian 5-year-olds with respect to the epistemic adverbs poate ‘maybe’ and sigur ‘certainly’. the paradigm is an improved version of the hidden object paradigm, where, instead of merely looking at an inaccessible entity, participants can now infer the presence of the entity on the basis of evidence (a shadow, as well as a specific sound). we argue that romanian children as young as 5 are able to derive implicatures with epistemic adverbs at an almost adultlike level. however, they exhibit the tendency to accept overly strong statements (i.e., statements where a certainty adverb is wrongly used instead of a possibility adverb) as optimal to a much higher degree than adults. this can be explained as a cognitive/communicative strategy to reduce multiple alternatives to a single one in cases of uncertainty. keywords. language acquisition; romanian l1; scalar implicatures; modality; epistemic adverbs; premature closure hypothesis 1. aim. the current paper is among the first studies to investigate the semantics and pragmatics of epistemic adverbs (possibility, not certainty implicatures) in child romanian (see also bleotu 2019). importantly, it employs a novel shadow play paradigm, an improved version of the traditional hidden object paradigm (hirst & weil 1982, noveck, ho & sera 1996, noveck 2001, ozturk & papafragou 2015, moscati, zhan & zhou 2017, a.o.), where participants have to make inferences about the presence of a hidden object/animal on the basis of evidence. while, in the traditional paradigm, participants have no direct access to the hidden object/animal and rely solely on reasoning, sometimes showing caution in their answers, in the shadow play paradigm, additional cues (the object/animal’s shadow and specific sound) support participants’ logical reasoning. as we will show, in the improved paradigm, children behave more adult-like than in the hidden object paradigm with respect to implicature-generation with the epistemic poate ‘maybe’. however, they behave non-adult-like in accepting as optimal overly strong statements, where certainty adverbs are wrongly employed instead of possibility adverbs. the paper is organized as follows: after presenting some background on the acquisition of epistemic modality in section 2, we present our novel shadow play paradigm in section 3. in section 4, we discuss the implications of our experimental results, and section 5 presents the conclusions of our research. * this research was supported by an xprag.de internship offered to adina camelia bleotu at zas berlin within the dfg project si games i: experimental game theory and scalar implicatures led by dr. anton benz (grant nr. be 4348/4-1). anton benz was supported by the bundesministerium für bildung und forschung (bmbf), grant nr. 01ug1411. nicole gotzner was supported by the dfg, grant nr. be 4348/4-2, through the priority program new pragmatic theories based on experimental evidence (spp 1727), and she is further supported by the dfg through the emmy noether programme (grant nr. go 3378/1-1). we are grateful to the students from the faculty of foreign languages and literatures, university of bucharest, who helped with the experiments, as well as to the director, tutors and children from no. 248 kindergarten, bucharest. authors: adina camelia bleotu, icub, university of bucharest (cameliableotu@gmail.com), anton benz, zas berlin (benz@leibniz-zas.de), & nicole gotzner, zas berlin (gotzner@leibniz-zas.de) and university of potsdam. proceedings of elm 1: 059-070, 2021 c©2021 adina camelia bleotu, anton benz and nicole gotzner published by the lsa with permission of the author(s) under a cc by license. 59 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2. background on the acquisition of epistemic modality. the experiment we conducted deals with epistemic modality, a type of modality which refers to the degree of speaker-commitment to the truth of the proposition expressed by the complement (kratzer 1981, coates 1983, sweetser 1982, 1990, palmer 1986, a.o.). epistemic modality translates into certainty or uncertainty about a situation, an attitude resulting from an inference based on evidence. this type of modality is usually set in contrast to root modality, which refers either to the necessity (obligation) or possibility (permission) of certain actions given society or moral rules (in the case of deontic modality) or to the ability or willingness (volition) to perform a certain action (in the case of dynamic modality). given the wide literature on the acquisition of epistemic modality, we will briefly touch on some essential aspects, and then focus on the experiments relevant for our study. in the acquisition literature, there is a general consensus that epistemic modality emerges after root modality (schatz & wilcox 1977, wells 1979, shepherd 1982, perkins 1983, stephany 1986). children produce deontic modals as early as 1;1, while they produce epistemic modals later (around 3). importantly, this order in acquisition is considered a reflection of a child’s mental development, that is, their theory of mind (wellman 1990, gopnik 1993, gopnik & wellman 1994, papafragou 1998, 2000). this refers to the ability to reflect upon one’s and others’ thoughts, to represent beliefs as belonging to someone else than themselves. interestingly though, not all epistemic modals behave in the same way: epistemic adverbs/adjectives (e.g., possible/ly, certain/ly) emerge before the age of 3 (o’neill & atance 2000, cournane 2015, veselinović and cournane 2020), whereas epistemic verbs (e.g., might, must) emerge after this age. this contrast in production is unexpected under an explanation that relies solely on theory of mind. thus, it has been proposed that, in addition to the cognitive account, there is a grammatical source for the acquisition of epistemic modal verbs (heizmann 2006, hacquard 2006, 2009, cournane 2015). in contrast to adverbs/adjectives, in the case of modal verbs, children have to figure out that the same items can express both root and epistemic meanings, or that modal verbs followed by lexical verbs in the progressive/perfective aspect usually give rise to epistemic meanings (kratzer 1981, papafragou 1998, 2000, avram 1999). the class of epistemic adverbs is not without its problems. interestingly, there seems to be a much higher number of epistemic possibility adverbs in comparison to epistemic necessity adverbs. this cannot be captured by theory of mind alone, but rather by frequency in the input (dieuleveut et al. 2019), or additional considerations, such as the necessity to resort to uncertainty markers in order to express uncertainty, but the absence of the necessity to resort to certainty markers in order to express certainty, as one can simply assert it. acquisition studies on epistemic modality have almost exclusively looked at epistemic verbs rather than adverbs (hirst & weil 1982, noveck, ho & sera 1996, noveck 2001, heizmann 2006, ozturk & papafragou 2015, moscati, zhan & zhou 2017, a.o.), but their methodology is essential for our purposes. all the experiments rely on some version of the hidden object paradigm, where participants have to infer the presence of a certain object/animal based on certain statements. however, there is variation in the conclusions about whether children behave adult-like or not. the first linguists to use this paradigm to investigate the semantics of epistemic modals were hirst & weil (1982), who implemented a look for the peanut task, where 3-to-6-year-olds children heard statements (with epistemic modals) about a peanut, and they had to find it. their results show sensitivity to modal strength: children looked more for the peanut in a certain location when the sentence they heard contained a strong (certainty) modal than when it contained a weak (uncertainty) modal. this shows children as young as 3 have a grasp of the modal scale. noveck, ho & sera (1996) and noveck (2001) further investigated the semantics and pragmatics of epistemic modal verbs through the box paradigm, a variant of the hidden object proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 60 https://doi.org/10.3765/elm https://www.elm-conference.net/ paradigm making use of boxes. for instance, participants were told that a covered box, box c has either the content of an open parrot + bear box (containing a parrot and a bear) or the content of an open parrot-only box (containing only a parrot). they were then asked to evaluate certain sentences describing the location of animals in box c. with respect to the semantics of epistemic modal verbs, noveck, ho & sera (1996) showed through a forced choice task that, unlike 5year-olds, 7-year-olds prefer weaker true statements over stronger false (overly strong) statements at an adult-like rate. in addition, noveck (2001) showed through a truth value judgment task that children derive implicatures with epistemic modals at a significantly lower rate than adults. however, these results may be considered slightly problematic given the complex nature of the task (the memory load, challenging instructions containing disjunctive statements, a.o.). ozturk & papafragou (2015) then investigated epistemic modals through a simplified version of the box paradigm, one where there were two boxes instead of three, and the instructions no longer contained disjunction as in noveck, ho & sera (1996) or noveck (2001). children as young as 4 were able to draw implicatures (though not at fully adult-like rates) in a forced choice task. moreover, children were also able to understand the meaning of epistemic modals in a variety of situations. however, they often had problems when faced with a situation open to multiple possibilities, and they had to evaluate an overly strong statement which referred to only one of these. similar conclusions have been reached by moscati, zhan & zhou (2017) on the basis of an eye-tracking experiment in the visual paradigm. children showed different fixation patterns than adults at the end of sentences containing the strong epistemic modal must in undetermined scenarios, which can be explained by their reducing multiple alternatives to a single one. while the results about how 5-year-olds treat overly strong epistemic statements seem to converge across experiments, the results about whether they derive scalar implicatures with epistemic modals do not. this could be an effect of the different tasks used, as children are known to perform more adult-like with felicity judgment/forced choice tasks than with truth value judgment tasks, but it could also be related to the fact that the hidden object paradigm may encourage participants to consider statements with weak epistemic modals optimal, given the fact that there is no direct access to the animal(s) hidden in the box, and children might be tempted to perceive the presence of an animal in the box as a possibility rather than a certainty. if this is so, however, it becomes unclear why children seem to accept overly strong statements: possible explanations involve cognitive considerations related to the reduction of uncertainty to certainty. in order to probe into this matter further, bleotu (2019) conducted an experimental study, investigating the semantic and pragmatic understanding of epistemic modals by romanian 5-yearolds, focusing on the more frequent adverbs sigur ‘certainly’ and poate ‘maybe’. the first experiment employed a coloring task where participants were asked to color certain drawings based on various statements containing (certainty/uncertainty) epistemic adverbs. in this experiment, children were sensitive to modal strength, always coloring the object in the color mentioned in the statement containing the certainty adverb, but only half of the time in the color mentioned in the statement containing the uncertainty adverb. the second experiment was a coloring version of the truth value judgment task in noveck (2001), aiming to see if children derive implicatures with epistemic adverbs. the results did not provide evidence that 5-year-olds derived implicatures, in fact, children derived no implicatures at all. interestingly, adults also showed rather low implicature rates (around 40%). overall, adults tended to give cautious answers, often rejecting pragmatically adequate sentences with sigur ‘certainly’, arguing that they could not be certain about the situation because they could not see what was going on with their own eyes. lack of direct access thus made adults hold back from asserting certainty. while noveck (2001), proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 61 https://doi.org/10.3765/elm https://www.elm-conference.net/ ozturk & papafragou (2015) and moscati, zhan & zhou (2017) report no such problems with adults, the problem noticed in bleotu (2019) calls for an improvement of the hidden object paradigm in such a way as to make sure that participants (adults and children) can embrace epistemic certainty even when they cannot access the hidden object. moreover, since both noveck (2001) and bleotu (2019) obtained low implicature rates with children with a truth value judgment task, a different task (e.g., a reward task) might produce more adult-like results. 3. the shadow play experiment. considering the indirect access challenge posed by the hidden object paradigm, we decided to modify the paradigm so as to make it easier for children (and adults) to reason about a hidden entity. in our paradigm, participants are supported in their inferences by extra cues such as the animals’ silhouettes and specific sounds. rationale and goal. the experiment we designed aims to investigate the mastery of the semantics and pragmatics of the epistemic modal adverbs sigur ‘certainly’ and poate ‘maybe’ by 5-year-olds. we decided to focus on epistemic adverbs rather than epistemic verbs, given the fact that trebuie ‘must’ with an epistemic meaning is very rare in adult romanian. using penncontroller (zehr & schwarz 2018), we implemented a novel shadow play paradigm, where participants can see the animals’ silhouettes and hear their specific sounds. we took inspiration from shadow play theatre, an ancient form of story-telling making use of shadows, as well as from by heizmann (2006), who used silhouettes behind a milky window to test whether german and english 3-year-olds are able to infer from a question such as who must be eating the banana? that the banana-eater is hidden from sight and cannot be right before their eyes. heizmann (2006) showed that, while the theory of mind definitely plays a part in the acquisition of modality, it seems that syntactic ambiguity or the lack thereof does too. while in contexts that are ambiguous between deontic and epistemic readings, children seem to prefer deontic readings over epistemic ones, children as young as 3 are able to understand epistemic verbs in a non-ambiguously epistemic context. the shadow play paradigm takes the idea of employing silhouettes as a starting point for testing the semantic and pragmatic understanding of epistemic adverbs. instead of focusing on indirect inferences (i.e., inferences about the lack of direct access to a certain entity), as in heizmann (2006), we focused on epistemic modal strength and scalar implicatures with epistemic adverbs. representing entities as silhouettes justifies the use of epistemic adverbs (given the lack of direct access to the animals). moreover, it makes the indirect evidence more ‘direct’, as participants are no longer in doubt about the animal’s location, but, instead, they can rely on evidence (the animal’s silhouette and specific sounds). in terms of task type, the paradigm asks participants to reward baby dragons with big or small apples depending on whether what they say is the best description of the situation or not. the best description task we used is a binary version of the ternary reward task from katsos & bishop (2011), where children reward statements with huge/big/small strawberries. we decided to make optimality rather than truth value (right/wrong) a reward criterion due to the higher number of scalar implicatures obtained previously with the best description task (bleotu, benz & gotzner 2020), as well as with the similar best response paradigm (gotzner & benz 2018). thus, given the use of the shadow play paradigm and of the best description task, we expect romanian 5-year-olds to perform somewhat better on the semantics and pragmatics of epistemic adverbs than in the previous tasks on modality (even if maybe not fully adult-like). participants. 35 5-year-olds (17 female and 18 male, age range: 5-6;6, mean age: 5;6) and a control group of 36 romanian adults (undergraduates from the faculty of foreign languages, at the university of bucharest) took part in the experiment in exchange for course credit. proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 62 https://doi.org/10.3765/elm https://www.elm-conference.net/ pretest. given the worry that children might not be able to understand the meaning of the best description in the instructions, we decided to run a superlative pretest before the experiment and filter out the children unable to handle superlatives. the pretest aimed to see how children understand superlatives both in deictic contexts, i.e., contexts where they could point to a certain object/animal, as well as in pragmatic contexts, i.e., contexts where children decided which statement best described a picture (see (1) and figure 1). (1) a. show me the tallest giraffe/a peach which is big, but not the biggest. b. this is a bear (informative)/an animal (underinformative)/a frog (false). figure 1: example pictures for the superlative pretest main experiment: methodology. in the main part of the experiment, participants are told there is a wizard who plays a shadow game with two baby dragons flurry and bindy. in the game, there are various animals who go and hide behind the curtain, but they come in front of the curtain one by one later on, at different stages. the baby dragons take turns to say who they believe the shadow belongs to, based on the evidence they have. participants have to reward them with a big apple if what they say is the best description of the situation and with a small apple otherwise. figure 2: the wizard and the baby dragons figure 3: the rewards: a big or a small apple the experimental materials involve several associated pictures and sentences referring to various groups of animals: a control/training group of two bunnies (orange and pink), and 4 testing groups of three animals of the same category (of different colors) each: dogs, frogs, cats, cows. the design of the pictures (see figures 4, 5, 6) tries to make it easy for participants to figure out the reference of the shadow (the main silhouette center-stage) and prevent processing difficulties (crain & thornton 1998), by presenting participants not only with information about the animals that are in front of the curtain (through a small image in the bottom part of the picture), but also with information about all the animals in the game (through a small image on the left). in total, participants saw 31 sentences (3 training sentences, 1x4=4 test sentences, 4x7= control sentences) containing poate ‘maybe’ or sigur ‘certainly’, presented in a randomized manner (see table 1). all the sentences (except for the practice ones) have the same structure: the epistemic adverb poate ‘maybe’ and sigur ‘certainly’ followed by the complementizer cǎ ‘that’ and an embedded sentence referring to the identity of the silhouette. importantly, optimal sentences with sigur ‘certainly’ (uttered by one dragon) are always followed by the corresponding proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 63 https://doi.org/10.3765/elm https://www.elm-conference.net/ underinformative sentences with poate ‘maybe’ (uttered by the other dragon). this contrast is maintained throughout to activate the modal scale and lead to implicature-generation. one in front scenario two in front scenario spossible1 optimal spossible2 optimal scertain1 overinfo scertain2 optimal spossible3 underinfo scertain3 false spossible4 false table 1: types of sentences tested per scenario the experiment comprises a training session and a testing session. in the training session, participants practice the reward task on a bunny shadow picture (see figure 4). subjects were presented with the sentences in (2): they were told which reward to choose (the small apple) in the first sentence, while, in the other sentences, they had to choose the reward themselves. figure 4: item for the training session (2) a. este un şoarece/o vacǎ. (false) ‘it is a mouse/a cow.’ b. este un iepuraş. (true/optimal) ‘it is a bunny.’ in scenario 1, the one in front scenario, one animal comes back in front of the curtain, in this case, the yellow dog (see figure 5). this allows us to test participants’ understanding that the situation has two possible outcomes: the silhouette belongs either to the red dog or the blue dog. subjects were presented with sentences such as those in (3): they had to choose a big apple for the optimal control statements in (3a) and a small apple for the overly strong statement in (3b). figure 5: one in front scenario figure 6: two in front scenario figure 7. disclosure (3) a. poate cǎ este cȃinele roşu/albastru. (optimal) ‘it is possible that it is the red/blue dog.’ b. sigur cǎ este cȃinele roşu. (overly strong) ‘it is certain that it is the red/blue dog.’ in scenario 2, the two in front scenario, two animals come back in front of the curtain (see figure 6). given such evidence, participants are supposed to reason that the silhouette can only belong to the blue dog. subjects were presented with the critical sentences (4a, b) and the proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 64 https://doi.org/10.3765/elm https://www.elm-conference.net/ control sentences (4c, d). subjects who only consider semantic meaning are expected to reward the baby dragon with a big apple in conditions (4a) and (4b), while subjects who strengthen the weak epistemic to ‘not certain’ are expected to give a small apple reward for (4a) but not for (4b). the scenario is thus critical for establishing whether subjects derive implicatures. (4) a. poate cǎ este cȃinele albastru. (underinfo) ‘it is possible that it is the blue dog.’ b. sigur cǎ este cȃinele albastru. (optimal) ‘it is certain that it is the blue dog.’ c. poate cǎ este cȃinele roşu. (false) ‘it is possible that it is the red dog.’ d. sigur cǎ este cȃinele galben. (false) ‘it is certain that it is the yellow dog.’ in the end, the identity of the animal is disclosed. while hirst & weil (1982) showed that disclosure does not affect experimental results, we decided to opt for disclosure regardless, as we felt it would keep subjects more engaged and quench their curiosity. results. we analyzed the accuracy of the answers in control statements, depending upon whether more than half of the answers were correct, and we excluded two adults from further analyses. in the case of children, we looked at the accuracy of the answers in the pretest and removed no participants, since all children gave more than 3 correct answers out of 6. moreover, since all children were more than half of the time accurate in the control sentences, they were all included in the analysis. the results were analyzed with logit mixed-effects models in r (2018). 3.5.1 whole data analysis. we computed a logit mixed-effects model with the factors group (adults, children), sentence type (underinformative, overly strong, control), as well as their interaction as fixed effects, and random by-item and by-participant slopes. the control statements of the adult group were chosen as the reference level. the results do not show significance for group, but there was a significant effect for sentence type (both for underinformative and overly strong statements), an interaction between group and underinformative sentences, as well as an interaction between group and overly strong sentences (see table 2). table 2: results of a glmer performed on the whole data 3.5.2 subset analysis. we divide the subset analysis into three parts: scalar implicatures, control statements, and overly strong statements. for scalar implicatures, the whole data analysis does not reveal precise information about implicature derivation since a participant’s choice of a small apple for underinformative sentences like (4a) simply indicates the degree to which participants rejected underinformative statements as not the best description of a situation. however, such rejection can happen for two reasons: (i) parameter estimate std. error z p intercept -0.036 0.098 -0.369 0.712 group (children) 0.174 0.142 1.217 0.224 sentence type overly strong 1.375 0.226 6.088 1.14e-09 *** sentence type underinformative 0.829 0.203 4.093 4.25e-05*** group (children): sentence type overly strong -1.904 0.296 -6.423 1.33e-10*** group (children): sentence type underinformative -0.996 0.277 -3.596 0.000323*** proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 65 https://doi.org/10.3765/elm https://www.elm-conference.net/ either because participants believe the stronger alternative is the best description, (ii) or because they believe the stronger alternative is false, hence, also not the best description. since we did not ask subjects to motivate their answers, the only way to determine their pattern of thinking is by looking at the corresponding stronger alternatives. on this basis, we can distinguish between pragmatic, logical, cautious and erroneous participants (see table 3). table 3: subjects’ patterns of responses in underinformative and optimal true sentences 66.18% of the answers produced by adults were scalar implicatures with epistemic adverbs: 22 adults consistently rejected underinformative statements (3 or 4 answers out of 4). 49.28% of children’s answers were scalar implicatures with epistemic adverbs: 15 children consistently produced implicatures (see figure 8). 49.28% of children’s answers were ‘overgenerous’ (logical): 14 were consistent in their answers. importantly, there were few erroneous or cautious answers, thus showing that the shadow play paradigm encourages logical reasoning. in the hidden object paradigm, subjects had to rely exclusively on logical reasoning to infer that sentences with certain are correct, and experimental results show that purely logical inference is too weak a basis for inferring certainty (bleotu 2019). the shadow play paradigm remedies such worries by providing additional clues that support (but do not replace) logical reasoning. figure 8: scalar implicatures per group in addition, we computed a logit mixed-effects model on a subset of the data, with the rate of scalar implicatures as the dependent variable, group as a fixed effect, and random by-item and byparticipant slopes. this model revealed no significant difference between the groups (β = −1.816, se = 1.0163, z = −1.787, p = 0.0739). as presented in table 4, children behaved adult-like with respect to the control sentences, as also revealed by running a logit mixed-effects model on the control data subset with group, truth and their interaction as factors, and item and participant as random effects. the results show nonsignificance for group (β = −0.578, se = 0.625, z = −0.925, p = 0.355) and the interaction between group and truth (β = −0.207, se = 0.519, z = −0.398, p = 0.691), but a significant truth effect (β = −2.1607, se = 0.391, z = −5.529, p < 0.001). most errors (93.91% errors for children, 87.5% errors for adults) were made in evaluating true sentences containing a possibility adverb, especially patterns of responses reward per statements two in front scenario underinformative (possible) informative (certain) adults children pragmatic logical cautious erroneous small apple big apple big apple small apple big apple big apple small apple small apple 66.18% 30.88 % 2.2% 0.74% 49.28% 49.28% 1.44% 0% proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 66 https://doi.org/10.3765/elm https://www.elm-conference.net/ in the one in front scenario, where participants had to acknowledge two alternatives. table 4: accuracy per group in control sentence as shown in figure 9, children tend to reward overly strong statements (using ‘certainly’ instead of ‘maybe’) with big apples to a much higher degree than adults (59.29% > 21.33%). the difference between groups is significant, as shown by a logit mixed-effects model on the overly strong data with group as a factor and random by-item and by-participant slopes (β = −4.158, se = 1.904, z = −2.184, p = 0.029). figure 9: yes to overly strong statements per group (with se) 4. discussion. our results show that romanian 5-year-olds are able to derive implicatures with epistemic adverbs in underinformative contexts. while a whole data analysis reveals a significant difference between children and adults, as in noveck (2001) or ozturk & papafragou (2015), the rejection of underinformative sentences does not equate with implicature-derivation. for this reason, we also performed several subset analyses taking into account adults’ answers to the strong alternatives of underinformative sentences. the subset analysis investigating implicaturegeneration showed children to be quite adult-like. this shows that children do not lack the capacity to derive implicatures, but, rather, they show sensitivity to the paradigm and task used. on the one hand, the current experiment makes use of an improved version of the hidden object paradigm used in previous experiments on epistemic modal items, namely, the shadow play paradigm, where children’s logical reasoning is supported by additional visual and acoustic cues. unlike in noveck (2001) or ozturk & papafragou (2015), children do not have to rely exclusively on logical reasoning in making inferences about a completely hidden animal, but they can, in addition, use silhouettes and sounds as further support. on the other hand, the current experiment also uses a different kind of task than the previous experiments on epistemic modals, namely, a binary reward task with optimality as a criterion. noveck (2001) employs a truth value judgment task on which 5-year-olds perform significantly different from adults. ozturk & papafragou (2015) employ a felicity judgment task on which children perform better than in noveck (2001), but still not adult-like. in contrast, in our experiment, children perform adult-like. since reward tasks are known to lead to more implicatures with quantifiers than other kinds of tasks (katsos & bishop 2011), it is not surprising that the same effect can be seen with epistemic adverbs. it is also accuracy per group in control sentences children adults optimal true control sentences false control sentences 75.2% 96.07% 82.59% 96.69% proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 67 https://doi.org/10.3765/elm https://www.elm-conference.net/ important that we were able to extend this result to romanian. improving the paradigm and changing the task reveal that romanian children as young as 5 have pragmatic abilities. in addition, we found that children also handle optimal true and false control statements with epistemic adverbs in an adult-like manner. thus, when children’s pragmatics is adult-like, the semantics seems to be in place also, an important result which is in line with the idea that semantic knowledge precedes pragmatic knowledge (noveck 2001). while adult-like in their derivation of scalar implicatures and their understanding of control statements, children are not adult-like in their treatment of overly strong statements, rewarding such statements with big apples to a higher degree than adults. the tendency to consider overly strong statements as adequate descriptions has previously been noticed in the studies by noveck, ho, sera (1996), ozturk & papafragou (2015), and moscati, zhan & zhou (2017). when faced with several possibilities, children tend to pick one possibility only. for example, in a situation where the silhouette could belong either to the blue dog or the red dog, there are children who give big apple rewards for statements which express certainty about the silhouette belonging to one of the dogs. there are several possible explanations for this. one possible explanation could be that children have a different semantics for the strong epistemic adverb sigur ‘certainly’: children could understand it as meaning ‘maybe’. nevertheless, this explanation is undermined by children’s correct assessment of control statements with sigur ‘certainly’, as well as by children’s sensitivity to epistemic adverb strength (see bleotu 2019). another possible explanation could be that children use a cognitive strategy to reduce uncertainty to certainty. such a hypothesis, also known as the premature closure hypothesis (acredolo & horobin 1987), argues that children’s answers reflect their cognitive intolerance of situations that allow multiple outcomes and their overarching preference for a single solution. another explanation could be that children’s answers reflect neither a faulty semantics for the strong epistemic adverb, nor a cognitive strategy to eliminate uncertainty, but rather a guessing communicative strategy, leading children to place a bet on one of the possibilities when the evidence is not conclusive. teasing apart the cognitive and the communicative account is difficult, especially if we embrace the view that communication mirrors cognition. children’s ‘guesses’ could be a reflex of a cognitive tendency to make certainty subjective. for instance, when evaluating the statement it is certain that it (the silhouette) is the red dog in a context where it is only possible that the silhouette is the red dog, one child explicitly motivated his big apple answer by saying that he likes red a lot. this suggests that ‘guesses’ might not be random, and, instead, cognitive/communicative reductions of uncertainty to certainty may be modulated by personal likes/dislikes. however, subjective reasons do not fully explain the results since children only ‘randomly’ choose after having already restricted the set of possible outcomes through an inference. importantly, when the silhouette can be the red or blue dog, children exclude the impossible outcome (giving small apples for false statements like it is certain that it is the yellow dog), and they infer the two possible alternatives (giving big apples for both optimal statements with possible). reducing uncertainty to certainty thus involves a logical step, where children infer the two possible outcomes, followed by a subjective step, where children make a choice between them, depending upon their own preferences and inclinations. 5. conclusion. the current experiment employed a novel shadow play paradigm in order to test romanian 5-year-olds’ semantic and pragmatic knowledge of epistemic adverbs. unlike more traditional versions of the hidden object paradigm, the shadow play paradigm gives participants additional evidence in order to help them perform in a more adult-like fashion. the results revealed children’s near adult-like ability to draw implicatures with epistemic adverbs. however, children seemed to accept overly strong sentences to a much higher degree than adults. this can be proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 68 https://doi.org/10.3765/elm https://www.elm-conference.net/ accounted for if one assumes a stage in language acquisition where, although children are able to infer that a situation has two outcomes, they make a subjective choice for one single outcome. references avram, larisa. 1999. auxiliaries and the structure of language. bucharest: unibuc acredolo, curt, & horobin, karen. 1987. development of relational reasoning and avoidance of premature closure. developmental psychology 23(1). 13–21. https://doi.org/10.1037/00121649.23.1.13. bleotu, adina c. 2019. what colouring can tell us about the acquisition of scalar items in child romanian. osf. september 29. osf.io/bwrvt. bleotu, adina c., anton benz, & nicole gotzner, n. 2020. where truth and optimality part. an experimental approach to epistemic adverbs. retrieved from osf.io/dh82r coates, jennifer. 1983. the semantics of modal auxiliaries. london: croom helm. cournane, ailís. 2015. revisiting the epistemic gap: evidence for a grammatical source. proceedings of the 39th annual to the boston university conference on language development (bucld39). 127-140. somerville, ma: cascadilla press. crain, stephen & rosalind thornton. 1998. investigations in universal grammar: a guide to experiments on the acquisition of syntax and semantics. cambridge, ma: mit press dieuleveut, anouk, annemarie van dooren, ailís cournane, & valentine hacquard. 2019. learning modal force: evidence from children’s production and input. in j. j. schlӧder, d. mchugh & f. roelofsen (eds.), proceedings of the 2019 amsterdam colloquium.111-122. gopnik, alison. 1993. how we know our minds: the illusion of first-person knowledge of intentionality. behavioral and brain sciences 16(1). 1–14. https://doi.org/10.1017/s0140525x00028636. gopnik, alison & henry m. wellman. 1994. the theory theory. in l. a. hirschfeld & s. a. gelman (eds.), mapping the mind: domain specificity in cognition and culture, 257–293. new york: cambridge university press. https://doi.org/10.1017/cbo9780511752902.011 gotzner, nicole & anton benz. 2018. the best response paradigm: a new approach to test implicatures of complex sentences. frontiers in communication 2(21). 1-13. https://doi.org/10.3389/fcomm.2017.00021. hacquard, valentine. 2006. aspects of modality. cambridge, ma: mit dissertation. hacquard, valentine. 2009. on the interaction of aspect and modal auxiliaries. linguistics and philosophy 32. 279-312. https://doi.org/10.1007/s10988-009-9061-6. heizmann, tanja. 2006. acquisition of deontic and epistemic readings of must and müssen. in tanja heizmann (ed.), university of massachusetts occasional papers in linguistics (umop) 34: current issues in language acquisition. amherst, ma: glsa, umass amherst. hirst, william & joyce weil. 1982. acquisition of epistemic and deontic meaning of modals. journal of child language, 9(3). 659–666. https://doi.org/10.1017/s0305000900004967. katsos, napoleon & dorothy bishop.2011. pragmatic tolerance: implications for the acquisition of informativeness and implicature. cognition 120 (1). 67-81. https://doi.org/10.1016/j.cognition.2011.02.015. kratzer, angelika. 1981. the notional category of modality. in h. j. eikmeyer & h. rieser (eds.), worlds, words, and contexts, 38-74. berlin: de gruyter noveck, ira. 2001.when children are more logical than adults. cognition 78(2). 165-188. https://doi.org/10.1016/s0010-0277(00)00114-1. noveck, ira a., simin ho & maria sera. 1996. children's understanding of epistemic modals. journal of child language 23(3). 621-643. https://doi.org/10.1017/s0305000900008977. proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 69 https://doi.org/10.3765/elm https://www.elm-conference.net/ o’neill, daniela k. & cristina m. atance. 2000. “maybe my daddy give me a big piano”: the development of children’s use of modals to express uncertainty. first language 20(58). 29– 52. https://doi.org/10.1177/014272370002005802 ozturk, ozge & anna papafragou. 2015. the acquisition of epistemic modality: from semantic meaning to pragmatic interpretation. language learning and development 11(3). 191-214. https://doi.org/10.1080/15475441.2014.905169. palmer, frank r. 1986. mood and modality. cambridge: cambridge university press. papafragou, anna. 1998. the acquisition of modality: implications for theories of semantic representation. mind & language 13(3). 370–399. https://doi.org/10.1111/14680017.00082. papafragou, anna. 2000. modality: issues at the semantics-pragmatics interface. oxford, england: elsevier. perkins, michael r. 1983. modal expressions in english. london: frances pinter. https://doi.org/10.1017/s000841310001094x. r core team. 2018. r: a language and environment for statistical computing. r foundation for statistical computing. vienna, austria. available online at https://www.r-project.org/. jeschull, liane & tom roeper. 2009. evidentiality vs. certainty: do children trust their minds more than their eyes? in crawford, j. (ed.), proceedings of the third conference on generative approaches to language acquisition north america (galana), 107–115. somerville, ma: cascadilla. shatz, marilyn & sharon a. wilcox. 1991. constraints on the acquisition of english modals. in a. gelman & j. p. byrnes (eds.), perspectives on language and thought: interrelations in development. 319–353. cambridge, ma: cambridge university press. https://doi.org/10.1017/cbo9780511983689.010. shepherd, susan.1982. from deontic to epistemic: an analysis of modals in the history of english, creoles, and language acquisition. in a. ahlqvist (ed.), papers from the 5th international conference on historical linguistics, 316-123. amsterdam: benjamins. https://doi.org/10.1075/cilt.21.36she. stephany, ursula. 1986. modality. in p. fletcher and m. garman (eds.), language acquisition: studies in first language development. london: cambridge university press. sweetser, eve. 1982. root and epistemic modals: causality in two worlds. berkeley linguistics society 8. 484–507. sweetser, eve. 1990. from etymology to pragmatics: metaphorical and cultural aspects of semantic structure. cambridge: cambridge university press. veselinović, dunja & ailís cournane. 2020. the grammatical source for missing epistemic meanings for modal verbs in child bcs. in t. ionin and j. e. macdonald (eds.), formal approaches to slavic linguistics (fasl 26), 417-436. ann arbor, mi: michigan slavic publications. wellman, henry. 1990. the child’s theory of mind. cambridge: mit press. wells, gordon. 1979. learning and using the auxiliary verb in english. in v. lee (ed.), cognitive development: language and thinking from birth to adolescence, 250–270. london: croom helm. zehr, jeremy & florian schwarz. 2018. penncontroller for internet based experiments (ibex). https://doi.org/10.17605/osf.io/md832. proceedings of elm 1: 059-070, 2021 adina camelia bleotu, anton benz and nicole gotzner: shadow playing with romanian 5-year-olds. epistemic adverbs are a kind of magic!. 70 https://doi.org/10.3765/elm https://www.elm-conference.net/ when transformer models are more compositional than humans: the case of the depth charge illusion dario paape* abstract. state-of-the-art transformer-based language models like gpt-3 are very good at generating syntactically well-formed and semantically plausible text. however, it is unclear to what extent these models encode the compositional rules of human language and to what extent their impressive performance is due to the use of relatively shallow heuristics, which have also been argued to be a factor in human language processing. one example is the so-called depth charge illusion, which occurs when a semantically complex, incongruous sentence like no head injury is too trivial to be ignored is assigned a plausible but not compositionally licensed meaning (don’t ignore head injuries, even if they appear to be trivial). i present an experiment that investigated how depth charge sentences are processed by transformer models, which are free of many human performance bottlenecks. the results are mixed: transformers do show evidence of non-compositionality in depth charge contexts, but also appear to be more compositional than humans in some respects. keywords. transformer, language model, depth charge illusion, compositionality 1. introduction. transformer-based language models like gpt-3 (brown et al. 2020) show impressive capabilities in terms of being able to produce naturalistic, usually grammatically correct, and often coherent text (e.g., dale 2021). rather than using recurrent and convolutional neural networks like previous state-of-the-art language models, the transformer architecture is based on a so-called self-attention mechanism: during training on large amounts of human-written text, the model learns which parts of the preceding context are important for predicting the next word (vaswani et al. 2017). using only this information, gpt-3 can “write” passable philosophical essays (elkins and chun 2020), and even scientific papers about itself (generative pretrained transformer et al. 2022). the outputs are often so convincing that humans cannot distinguish between machine-generated and human-written text (clark et al. 2021, uchendu et al. 2021). at the same time, however, it is relatively easy to unmask transformer models if one knows what to look for. for instance, due to their architecture and training regime, transformers often fail at simple arithmetic (floridi and chiriatti 2020, patel, bhattamishra and goyal 2021), arrive at bizarre deductions in scenarios that require real-world knowledge, and sometimes output obvious non sequiturs with sudden and extreme topic shifts that would be absurd coming from a human writer or speaker (marcus and davis 2020). furthermore, given that any “knowledge” about the world that may be encoded in the model is not grounded in experience or reasoning but is filtered through language and its statistical properties (e.g., alberts 2022), such as the frequent co-occurrence of certain terms, transformers often resort to heuristics: they produce associatively plausible rather than factually correct answers to information questions (sobieszek and price 2022), and to some extent rely on simple lexical overlap between a premise and a hypothesis to predict entailment or non-entailment (mccoy, pavlick and linzen 2019). *the author would like to thank the vasishth lab members, yuhan zhang, and lisa levinson for helpful comments and suggestions. author: dario paape, department of linguistics, university of potsdam (paape@uni-potsdam.de). proceedings of elm 2: 202-218, 2023 c©2023 dario paape published by the lsa with permission of the author(s) under a cc by license. 202 https://doi.org/10.3765/elm https://www.elm-conference.net/ that the “knowledge” encoded by transformers is often heuristic in nature is also highlighted by so-called mispriming effects: for instance, bert (devlin et al. 2018), a close cousin of gpt-3, when asked to complete the input talk? birds can . . . , will produce talk as the most likely continuation, but will produce fly as the most likely continuation when the prompt is birds cannot . . . (kassner and schütze 2019). similarly, gpt-3 will assume that a mixture of cranberry juice and grape juice is poisonous if the linguistic context suggests that dangerous substances are being mixed (marcus and davis 2020). even though several of these limitations can be overcome by targeted training and/or the addition of symbolic knowledge (see helwe et al. 2021 for a review), it is nevertheless striking that untargeted training on very large language corpora often results in reliance on relatively shallow processing strategies. the implicit or explicit gold standard against which language models are usually compared and evaluated in terms of their “shallowness” is human performance. however, there are many tasks related to language and reasoning on which humans do not perform well, which casts doubt on this rationale (linzen and baroni 2021). many failures of human reasoning can also be seen as being the result of mispriming effects: for instance, the cognitive reflection test (frederick 2005) and its extension (thomson and oppenheimer 2016) contain questions such as how many cubic feet of dirt are there in a hole that is 3’ deep x 3’ wide x 3’ long?, to which 84% of participants answer “27” despite the correct answer being “none”. relatedly, when asked how many animals of each kind did moses take on the ark?, 81% of participants answer “two” even after having been instructed to look out for possible errors in the question, and even though they know that the biblical story is about noah (erickson and mattson 1981). human blindness to incongruous information in otherwise highly congruent contexts and the tendency to fall for verbal misdirection generalize to examples from different thematic domains (e.g., barton and sanford 1993, cook et al. 2018), and are not reducible to the default assumption that interlocutors always produce sensible statements and requests (reder and kusbit 1991). in light of results such as these, it has been proposed that human language processing is partly heuristic and often just “good enough” (e.g., ferreira and patson 2007, christianson 2016): instead of constructing detailed syntactic and semantic structures based on compositional rules, people may sometimes use high-level language statistics and world knowledge to derive a “quick and dirty” approximation of meaning. as a case in point, the meaning of implausible passive sentences such as the dog was bitten by the man is often converted into that of an active sentence with reversed roles (the dog bit the man), presumably because this meaning is a priori more plausible, and because the agent of an event is usually mentioned first in english (ferreira 2003, christianson, luke and ferreira 2010). 1.1. the depth charge illusion. the difference between heuristic language processing and “reasoning” in transformers and in humans is that the human version can usually be neutralized by explicitly pointing out the problematic element(s) in a given sentence (e.g., erickson and mattson 1981, barton and sanford 1993), or by explaining the invalidity of a given inference and explaining the correct solution (e.g., van benthem 2008, claidière, trouche and mercier 2017, calvillo, bratton, velazquez, smelter and crum 2022). however, the incorrect “solutions” to some reasoning problems, like the monty hall problem (e.g., vos savant 1997, rosenthal 2008) and the wason selection task (wason 1968), famously tend to resist being explained away in this manner. proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 203 https://doi.org/10.3765/elm https://www.elm-conference.net/ in the linguistic domain, an example that also has the property of persistence in the face of corrective explanation is the so-called depth charge illusion, which was first discussed by wason and reich (1979). the depth charge illusion occurs in (1), which is often interpreted to mean don’t ignore head injuries, even if they appear to be trivial. (1) no head injury is too trivial to be ignored. anecdotally, a small subset of people immediately recognizes that this sentence is incorrect, a second subset is open to considering the possibility that it could be incorrect but cannot see why, and a third subset will stubbornly insist that it is correct, even after being confronted with the following argumentation: 1. the phrase too trivial to be ignored is semantically incongruous, because it presupposes that something can be so trivial that it should not be ignored (compare x is too young to die, which translates to x is so young that they should not die). 2. the incongruity is not removed by the initial negation: asserting that no head injury has the incongruous property of being too trivial to be ignored does not make the property itself any less incongruous. 3. the initial negation does cause the overall statement to be affirmative, contrary to the plausible misinterpretation (don’t ignore head injuries): in abstract terms, if no x is too y to be z’ed, this means that no x crosses the threshold beyond which it should not be z’ed, meaning that all x should be z’ed (that is, ignored). 4. compositionally, the sentence thus means ignore all head injuries, even if they appear to be trivial. the sentence can be made compositionally sensible by changing it to no head injury is too trivial to be noticed/treated or to no head injury is trivial enough to be ignored, but speakers will often find these variants difficult to process or even reject them as being malformed. despite broad agreement in the literature that the “don’t ignore” interpretation of (1) is not compositional, not all scholars agree that the depth charge illusion is due to a processing error, partly because it is so persistent. explanations generally fall into three categories: the classic shallow processing account (wason and reich 1979, paape, vasishth and von der malsburg 2020), an account based on the alleged idiomaticity of the construction no x is too y to z (cook and stevenson 2010, fortuin 2014), and an account based on unconscious correction of an assumed speech error (zhang, ryskin and gibson 2022). the shallow processing account claims that due to the syntactic and semantic complexity of (1), working memory becomes overloaded at some point and compositional processing breaks down or is suspended, presumably when the implicit negation contained in the word too is combined with the initial negation (paape et al. 2020). readers then use their world knowledge in combination with superficial language heuristics (duplex negatio affirmat; no head injury is too trivial . . . → all head injuries are too dangerous . . . ) to derive a plausible meaning. by contrast, the idiomatic or construction-based account assumes no breakdown. instead, its proponents claim that the no x is too y to z construction is a stored grammatical unit that can, by virtue of its idiomaticity, violate the compositionality principle and be “legally” interpreted to mean no x should be z’ed. finally, the error-correction account claims that readers combine prior expecproceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 204 https://doi.org/10.3765/elm https://www.elm-conference.net/ tations about plausible utterances with expectations about plausible speech errors to reconstruct the presumably intended sentence (no head injury is so trivial as to be ignored) and derive its meaning. the proposal that the human sentence processor uses prior expectations about sentence meanings and speech errors to “repair” potentially degraded input has also been applied to other cases of linguistic illusions (e.g., gibson, bergen and piantadosi 2013, frazier and clifton 2015). all proposed accounts have empirical weaknesses: shallow processing cannot explain why readers usually don’t consciously notice the complexity overload that causes compositional processing to break down, or why the illusion often cannot be explained away: if complexity is the problem, taking one’s time and working out the compositional meaning step by step should always lead to success. conversely, the construction-based account cannot explain why the illusion can sometimes be explained away, why some people appear to be immune to it, and why it generalizes to distinct but compositionally similar constructions (e.g., too . . . as that in german; paape et al. 2020). finally, the bayesian error-correction account cannot explain why readers usually cannot consciously access and report the assumed error correction (“i believe the speaker/writer made a mistake here”) even after multiple passes over the sentence,1 and why putting the incongruity in focus by changing the word order weakens the illusion (too trivial to be ignored is surely no head injury in german; paape 2021). investigating how transformers handle depth charge sentences may provide a way out of the empirical conundrum. unlike humans, transformers don’t have limited working memory capacity, so they don’t experience complexity overload. humans suffer from the “now or never” bottleneck (christiansen and chater 2016), that is, they must quickly and incrementally integrate incoming information before it is forgotten. transformers, by contrast, don’t process sentences incrementally but holistically, that is, they always have full access to all words in the sentence unless this access is deliberately limited (kahardipraja, madureira and schlangen 2021). on the other hand, transformers are prone to learning heuristics rather than compositional rules (e.g., mccoy et al. 2019), and are known struggle with negation (kassner and schütze 2019, hossain, kovatchev, dutta, kao, wei and blanco 2020, hosseini, reddy, bahdanau, hjelm, sordoni and courville 2021), so that they might to some extent mimic human “good enough” processing of depth charge sentences. by contrast, under the construction-based account, the transformer would need to learn from the training data that the no x is too y to z construction cannot only be used compositionally (no head injury is too trivial to be treated) but also non-compositionally, that is, idiomatically. however, the non-compositional variant is relatively rare: cook and stevenson (2010) report 170 instances of the construction in a written corpus of 1.1 billion words, of which 80% were compositional (e.g., no risk [is] too small to eliminate). fortuin (2014) reports only 13 instances of the “negative” (idiomatic) construction in a corpus of similar size. transformers tend to overgeneralize when the number of exceptions to a compositional rule is small, showing a) that compositionality is learned to some degree and b) that a certain amount of counterevidence is needed to “memorize” exceptions (hupkes et al. 2020).2 1even in an experimental context where the task is to correct incorrect use of too and enough, participants propose corrections to depth charge sentences only in about 30% of trials, and in most cases suggest changing the adjective (e.g., trivial → dangerous; o’connor 2015), which leaves the compositional implausbility of the overall interpretation (“don’t ignore head injuries”) intact. 2interestingly, even if the depth charge illusion is best characterized as a processing error in humans, if the error proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 205 https://doi.org/10.3765/elm https://www.elm-conference.net/ regarding the bayesian error-correction account of zhang et al. (2022), it is not entirely clear how it could be applied to transformers: transformers don’t have “common sense” learned from the real world against which they can evaluate sentence meanings, nor do they have a notion of speech errors, much less a subconscious error correction mechanism. to a standard transformer, there is no “noise” — every sentence seen during training is data, and the model can only adapt by allowing for more variability in its parameters (michel and neubig 2018, passban, saladi and liu 2020). there may be some notion of a “plausible utterance” encoded in the model, in the sense that words with certain meanings tend to co-occur, possibly even in syntactically similar environments, but it is unlikely that the model also encodes the assumption that unintended errors can lead to malformed sentences, and is able to reconstruct the original meaning. in what follows, i present an experiment in which different transformer models were tested on the no x is too y to z construction and a variety of control constructions. the experiment is exploratory in nature: the aim was not to conclusively answer the question of how the depth charge illusion arises in humans, but to see whether a system trained on many terabytes of text, but without incremental processing, memory bottlenecks, or error correction mechanisms would show the illusion or not. 2. experimental study. the purpose of the study was to assess the probability different transformer models assign to the word ignored as the next word after seeing the preamble no head injury is too trivial to be . . . . the log probability of ignored is treated as the dependent variable, and higher probabilities are taken to indicate a stronger depth charge illusion. this is a simplification, as the models may also assign high probabilities to continuations such as overlooked or forgotten about, which are semantically similar to ignored. to solve this problem, one can ask human coders to classify the continuations into “ignore-like” and “treat-like” categories, and then sum the relevant probabilities. however, there is some uncertainty as to which coding scheme should be used, as the relevant dimension of semantic similarity is not easy to capture (paape et al. 2020, o’connor 2015). i thus restrict my analysis to the single token ignored here. in sections 2.4 and 2.5, the overall distribution of continuations is discussed in more detail. 2.1. materials. the experiment had 9 conditions overall, as shown in (2). here, compositional is used as shorthand for “has a sensible meaning under a compositional analysis”, whereas not compositional is used as shorthand for “does not have a sensible meaning under a compositional analysis”. conditions (2-a), (2-b) and (2-c) are the depth charge conditions, while the rest are control conditions designed to test whether the models have encoded knowledge about negation, scales, and the degree particles too and so. the control conditions are syntactically and/or semantically less complex than the depth charge conditions, and a transformer that has encoded the relevant knowledge should consistently assign higher probability to ignored in the compositional conditions compared to the non-compositional conditions. is frequent enough, a transformer would presumably treat it as evidence of a grammaticalized rule exception. in addition, there are several academic papers on the depth charge illusion (see paape 2021 for a review), in addition to discussions on several public web forums (e.g., https://english.stackexchange.com/questions/ 91612/no-head-injury-is-too-trivial-to-ignore), which may become part of the training data of transformer models and provide “evidence” of the construction being used. proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 206 https://doi.org/10.3765/elm https://www.elm-conference.net/ (2) a. no head injury is trivial enough to be → ignored (compositional) v b. no head injury is too trivial to be → ignored (not compositional) x c. some head injuries are too trivial to be → ignored (not compositional) x d. no head injury is so trivial as to be → ignored (compositional) v e. no head injury is so trivial as to not be → ignored (not compositional) x f. head injuries that are too trivial will be → ignored (compositional) v g. head injuries that are not too trivial will be → ignored (not compositional) x h. head injuries that are trivial are more likely to be → ignored (compositional) v i. head injuries that are trivial are less likely to be → ignored (not compositional) x replacing too in the depth charge sentence (2-b) with enough in (2-a) yields a compositionally sensible meaning when the sentence is completed with the verb ignored. humans do indeed produce compositional, “ignore-like” completions for this construction in the majority of trials (o’connor 2015). the third depth charge condition (2-c) with some instead of no also leads to more compositional completions in humans, but in this case the completions are “treat-like” rather than “ignorelike” (paape et al. 2020). humans also assign lower sensibleness ratings to some-sentences ending with ignored, suggesting that the initial negation is crucially involved in “masking” the incongruity of the degree phrase in (2-b) and creating the depth charge illusion (paape et al. 2020, paape 2021). to see if differences between conditions generalize across different sentence contexts, the 32 german depth charge items used by paape et al. (2020) were translated into english and adapted to fit the design shown in (2). only the versions with negative adjectives were used (e.g., no plan is too unrealistic to be → scrapped, no physical theory is too implausible to be → dismissed). across all sentences, the dependent variable was the log probability of the verb used in the paape et al. rating experiments. 2.2. tested models. four models were tested. the first two were models of different sizes from the gpt-3 family: ada, the least powerful version of gpt-3, which “can perform tasks like parsing text, address correction and certain kinds of classification tasks that don’t require too much nuance” and davinci, the most powerful version, which “shines [...] in understanding the intent of text” and “is quite good at solving many kinds of logic problems” (https://beta.openai.com/docs/models/gpt-3). the third model under consideration was jurassic-1-jumbo, released by ai21 labs, which is similar in size to davinci but outperforms it in terms of predictive accuracy on many corpora (lieber et al. 2021). the fourth model was roberta, a retrained version of bert with improved performance (liu et al. 2019; https://huggingface.co/roberta-large). roberta’s training regime is somewhat different from that of gpt-3 and jurassic-1: like bert, roberta is bidirectional, that is, it not only considers the context to the left of a word but also the context to the right. roberta has 355 million parameters, and is thus similar in size to gpt-3 ada3. gpt-3 davinci and jurassic-1jumbo are much larger, with about 175 billion parameters each. the gpt-3 models were queried via the openai api, jurassic-1-jumbo was queried via the 3this assumes that ada corresponds to the 350m version of gpt-3 reported by brown et al. (2020) (https: //blog.eleuther.ai/gpt3-model-sizes/). proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 207 https://doi.org/10.3765/elm https://www.elm-conference.net/ -12 -10 -8 -6 (a) no ... enough (b) no ... too (c) so me ... too (d) so ... as to (e) so ... as to not (f) are too (g) are not to o (h) m ore lik ely (i) less l ikely lo gp ro b (i gn or ed ) model gpt-3 ada gpt-3 davinci jurassic-1-jumbo roberta figure 1: log probability of the critical verb (e.g., ignored) by construction and model. error bars show 95% confidence intervals across all 32 items. ai21 labs api, and roberta was queried via the huggingface api (wolf et al. 2020). for multi-token completions (e.g., . . . to be ruled out) the log probabilities of the generated tokens were summed. the pretrained models were used as-is; no fine-tuning of any kind was carried out. 2.3. results and bayes factor analysis. figure 1 shows the results by model and condition. in order to gauge the amount of statistical evidence in the data for differences between conditions and between models, the returned log probabilities were analyzed using a linear mixedeffects model (lmm) in stan (stan development team 2022) via the brms package (bürkner 2017) in r (r core team 2022). the lmm assumed a gaussian likelihood, and contained a fixed effect of model, which was treatment-coded with gpt-3 ada as the baseline, as well as random intercepts by sentence and random slopes for model by sentence. for the conditions, sum contrasts were defined in the following way: • (2-b) versus (2-a): too −1 versus enough +1 (compositional versus non-compositional) • (2-c) versus (2-b): some −1 versus no +1 (negation masks incongruity) • (2-d) versus (2-e): as to not −1 versus as to +1 • (2-f) versus (2-g): are not too −1 versus are too +1 • (2-h) versus (2-i): less likely −1 versus more likely +1 normal(0,1) priors were used for all contrasts, and bayes factors were computed using the bayestestr package (makowski, ben-shachar and lüdecke 2019). the model code and the data are available at https://osf.io/yw8vk. bayes factors are interpreted according to the scheme given by andraszewicz et al. (2015). the results of the analysis are as follows (bf is bf10):4 4because i am only interested in evidence for differences between models, i do not distinguish cases in which the null hypothesis is favored, that is, in which there is evidence in favor of the models behaving similarly, from cases in which the evidence doesn’t favor either hypothesis. proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 208 https://doi.org/10.3765/elm https://www.elm-conference.net/ • for gpt-3 ada as the baseline, strong evidence for positive differences for all contrasts (all bfs > 15), except for the too / enough contrast (bf = 1.6). • too / enough contrast: moderate evidence for larger difference in jurassic-1-jumbo than in gpt-3 ada (bf = 5.3). • some / no contrast: no evidence that other models perform differently from gpt-3 ada (all bfs < 1.1). • as to not / as to contrast: moderate evidence for larger difference in roberta than in gpt-3 ada (bf = 6.7). • are not too / are too contrast: strong evidence for larger difference in gpt-3 davinci than in gpt-3 ada (bf = 27), moderate evidence for larger difference in jurassic-1-jumbo than in gpt-3 ada (bf = 3.2). • less likely / more likely contrast: no evidence that other models perform differently from gpt-3 ada (all bfs < 1.7). 2.4. example completions by model. example sentence completions produced by the different models are shown below, along with indicators showing whether the completion is compositionally sensible or not. the prompt is underlined in each example. (3) gpt-3 ada a. no head injury is too trivial to be counted as a crime. (compositional) v b. some head injuries are too trivial to be taken lightly. (not compositional) x c. head injuries that are trivial are more likely to be fatal. x d. head injuries that are trivial are less likely to be fatal. v (4) gpt-3 davinci a. no head injury is too trivial to be ignored. any recent head injury, no matter how minor, should be included in the patient’s history. x b. some head injuries are too trivial to be treated, dr. benson acknowledged. v (5) jurassic-1-jumbo a. no head injury is too trivial to be noticed by a parent. v b. no head injury is too trivial to be ignored. all head injuries need to be taken seriously. x (6) roberta a. head injuries that are too trivial will be punished. ?? b. some head injuries are too trivial to be ignored. x the completions are interesting in multiple regards. completions like (3-a), (4-b) and (5-a) appear to be compositional, that is, there is no depth charge illusion in these examples. on the other hand, examples (4-a) and (5-b) show that when the models produce non-compositional completions, they will occasionally follow them up with semantically matching continuations. the pair (3-c)/(3-d) shows that gpt3 ada’s performance on the relatively straightforward control conditions is far from perfect: it is highly unlikely that humans would ever produce fatal in (3-c). completions (3-b) and (6-b) show that transformers produce non-compositional completions in the some condition as proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 209 https://doi.org/10.3765/elm https://www.elm-conference.net/ well, which happens very rarely in humans (paape et al. 2020). completion (6-a) by roberta is completely unexpected, unless it is taken to mean that whoever caused the head injury is going to be punished, in which case it would be a non-compositional completion. 2.5. forward and backward mask-filling with roberta. as mentioned above, roberta is bidirectional, and can thus produce completions based on rightward as well as on leftward context, that is, roberta is able to “retrodict” words based on future input. this is achieved by putting a [mask] token in place of the to-be-inserted word. the masking feature can be used to take a closer look at compositionality in roberta: given an input string like no head injury is too [mask] to be ignored, does roberta produce an adjective that results in a compositionally well-formed degree phrase? it is also worthwhile to look at the distribution of completions: even if the most likely completion is compositional, there may be an alternative, non-compositional completion with a similarly high probability, or the other way around. roberta’s top 4 verb and adjective completions with their associated probabilities for a set of example sentences are shown in (7) below. (7) a1. no head injury is too trivial to be [mask] addressed – 14%, treated – 9%, considered – 7%, ignored – 3% a2. no head injury is too [mask] to be ignored serious – 36%, minor – 10%, severe – 8%, small – 8% b1. no potential habitat is too nutrient-poor to be [mask] developed – 17%, exploited – 17%, explored – 11%, viable – 3% b2. no potential habitat is too [mask] to be ruled out small – 24%, remote – 12%, obscure – 4%, good – 4% c1. no chapter is too irrelevant to be [mask] archived – 6%, read – 5%, included – 3%, forgotten – 3% c2. no chapter is too [mask] to be skipped important – 50%, long – 13%, short – 7%, boring – 5% d1. no physical theory is too implausible to be [mask] tested – 40%, proven – 5%, tried – 4%, considered – 3% d2. no physical theory is too [mask] to be dismissed old – 11%, weak – 9%, absurd – 6%, good – 5% for the selected examples, most of roberta’s completions are compositional, though there are also non-compositional completions with relatively high probabilities. the picture differs markedly from human performance: humans complete no head injury is too trivial to be . . . and no chapter is too irrelevant to be . . . with “ignore-like” and “skip-like” continuations about 80% of the time (paape et al. 2020, appendix b). on the other hand, compositionally “retrodicting” the adjective seems to be difficult for roberta in some contexts: roberta mostly generates noncompositional adjective completions in (7-b2) and (7-d2), even though masking the verb mostly results in compositional completions for the same sentences in (7-b1) and (7-d1). for (7-a2) and (7-c2), on the other hand, the adjective completions are mostly compositional, just like the verb proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 210 https://doi.org/10.3765/elm https://www.elm-conference.net/ completions in (7-a1) and (7-c1). overall, roberta’s performance shows quite a lot of variability between sentences, which matches what is observed in humans (paape et al. 2020, paape 2021, zhang et al. 2022). possible sources of this variability are discussed below. 3. discussion. the aim of this paper was to investigate whether transformer-based language models show the depth charge illusion, in which a compositionally incongruous sentence (no head injury is too trivial to be ignored) is given an unlicensed but plausible interpretation (don’t ignore head injuries, even if they appear to be trivial). transformers are an interesting test case, given the range of proposed explanations for the illusion in human readers: theoretical proposals range from processing breakdown and recovery through world knowledge and superficial language heuristics (wason and reich 1979, paape et al. 2020), to the existence of an idiomatic no x is too y to z construction (cook and stevenson 2010, fortuin 2014), to bayesian speech error correction (zhang et al. 2022). transformers don’t experience processing breakdown, may not have seen enough instances of the hypothesized no x is too y to z construction to encode it, and have no means of distinguishing between “normal” training data and speech errors. the experimental results yielded some evidence that the depth charge illusion is present in transformers of different types and sizes: across 32 test sentences, all considered models (gpt-3 ada, gpt-3 davinci, jurassic-1-jumbo, and roberta) assigned higher probabilities to completions like ignored when the sentence began with a negation (no head injury . . . ) compared to when it did not (some head injuries . . . ), even though the completion results in an internally incongruous degree phrase (too trivial to be ignored) in both cases. furthermore, apart from jurassic-1-jumbo, none of the models appeared to distinguish between too and enough in negated contexts, even though the two degree particles have opposite meanings. at the same time, however, the transformer models showed evidence of compositional processing in control contexts such as head injuries that are trivial are less likely to be . . . [*ignored], suggesting that the required syntactic and semantic rules have been encoded. taken at face value, these results suggest largely parallel effects between human readers and transformers with regard to the depth charge illusion, despite the presumably very different underlying processing mechanisms. however, a closer look at the transformers’ sentence completions revealed that they produce a variety of un-humanlike continuations for the control conditions, suggesting that their grammatical “knowledge” may not be as deep as the high-level results suggest (bender et al. 2021). on the other hand, the mask-filling patterns of roberta suggested that roberta is often more compositional than human readers: for a variety of test sentences, including the most famous example no head injury is too trivial to be . . . , roberta showed a preference for compositional completions like addressed, unlike human participants (paape et al. 2020, o’connor 2015). at the same time, non-compositional completions did also appear in the list of most likely tokens, and even dominated for some sentences, especially in “retrodictive” contexts, that is, when roberta had to fill in the adjective based on the verb (e.g., no potential habitat is too [small] to be ruled out). what are the implications of these findings for the empirical deadlock between the competing psycholinguistic accounts of the depth charge illusion? proponents of the construction-based view could argue that the transformer models have picked up on the no x is too y to z construction to some extent, but haven’t fully mastered it yet, presumably because they haven’t encounproceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 211 https://doi.org/10.3765/elm https://www.elm-conference.net/ tered enough instances of it in the input data, and because the construction is arguably ambiguous between a compositional and a non-compositional version (cook and stevenson 2010, fortuin 2014). a preference for the compositional reading wouldn’t be surprising under this view, given that transformers are known to struggle with infrequent and/or idiomatic constructions (hupkes et al. 2020, dankers et al. 2022). this suggests some promising avenues for future research: providing disambiguating context should allow the transformer to identify the intended reading, as it arguably does for humans (fortuin 2014), and additional training on the construction should increase accuracy, in the sense that contextual cues should be more reliably identified. meanwhile, proponents of the superficial processing account could argue that even though transformers don’t experience processing breakdown, they could nevertheless be using heuristics to process depth charge sentences. this is a plausible assumption, given that transformers are known to learn heuristics in other settings, including negation processing (mccoy, pavlick and linzen 2019, helwe, clavel and suchanek 2021). that the models haven’t achieved human-like syntactic and semantic competence in terms of processing scales (more trivial → higher probability of ignoring), degrees (too trivial) and negation is clear from the many examples in which they produced compositionally incongruous completions in the control conditions. at the same time, however, the models do not appear to use the simplest possible heuristic for dealing with depth charge sentences: to ignore the beginning of the sentence and only locally evaluate the degree phrase too trivial to be ignored, which is always incongruous, irrespective of whether the sentence begins with no or with some. this strategy is unlikely to be used by human readers, who are limited by their incremental left-to-right processing and the “now or never” bottleneck (christiansen and chater 2016), which may lead them to incorrectly combine no and too before they even reach the verb (paape et al. 2020). transformers, on the other hand, do not have this limitation, and yet they have apparently learned to pay attention to the initial no in depth charge contexts. where does this leave the superficial processing account? the following scenario is possible: even larger transformer models with even more (or better) training may eventually acquire human-like syntactic and semantic competence, but may not exhibit the depth charge illusion, because their output is not limited by performance factors.5 resistance to the illusion may also gradually increase with scale and training, as the amount of abstract compositional knowledge and, presumably, transfer ability in the system increases. larger models typically do perform better on language tasks, though the current data do not show evidence of scale effects: the strength of the depth charge illusion was similar across models of very different sizes, though jurassic-1-jumbo showed some evidence of distinguishing more between too and enough than the other models. proponents of the bayesian error-correction account could argue that despite the absence of “reasoning” in transformers, the models may have learned to correct for speech errors to some extent. transformer-based language models often acquire unexpected capabilities that they were not explicitly trained for (e.g., radford et al. 2019, brown et al. 2020), and deep learning has been touted as a potentially powerful approach to grammar correction (dale and viethen 2021), so error 5this should also extend to other linguistic illusions that have been attributed to performance factors, such as agreement attraction (the key to the cabinets are on the table). however, the precise way in which transformers encode syntactic structure and linguistic dependencies may be inherently error-prone (finlayson et al. 2021, ryu and lewis 2021), so that complete immunity may be impossible. proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 212 https://doi.org/10.3765/elm https://www.elm-conference.net/ repair may be a latent capability of such models. however, reasoning about what the originally intended form or message of a given linguistic utterance was (no head injury is so trivial as to be ignored; zhang et al. 2022) and how it might have been transformed into its observed form (no head injury is too trivial to be ignored) is a very complex task. this type of reasoning would presumably require some approximation of a theory of mind, as well as an approximate model of human speech production, relevant topic knowledge (should all head injuries be treated?), and pragmatics. capabilities that resemble common-sense reasoning may be present in current transformers to some extent, but, as kejriwal et al. (2022) have recently argued, a more diverse set of empirical tools is needed to find out how much “common sense” is really encoded in the models. a deeper understanding of how transformers process the depth charge construction will, in all likelihood, benefit future research into why most humans struggle with the construction, and why there are such large differences between people and between specific sentences. at the sentence level, the interpretation of depth charge stimuli partly depends on the strength of world knowledge associated with the sentence (paape et al. 2020, zhang et al. 2022), as well as its sentiment polarity, semantic cohesion, and word order (paape 2021). investigating whether transformers are also sensitive to these factors would yield further insights into how similar the mechanisms behind the illusion are in transformers and humans. at the participant level, working memory capacity has been investigated as a potential source of variability, yielding a null result (paape et al. 2020). a more promising factor may be language experience. an interesting approach would be to try and fine-tune transformers in such a way that they become immune to the depth charge illusion, and to then see if a similar approach can be used to immunize human participants. “natural” immunity appears to be rare in humans but may also depend on language experience. if immunization is possible, this would support the original intuition of wason and reich (1979) that the depth charge illusion is a processing error, and weaken theories that see the illusion as a result of grammaticalization or pragmatic reasoning. references alberts, l., 2022. is it possible not to cheat on the turing test? exploring the potential and challenges for true natural language ‘understanding’ by computers. arxiv preprint arxiv:2206.14672 doi:https://doi.org/10.48550/arxiv.2206.14672. andraszewicz, s., scheibehenne, b., rieskamp, j., grasman, r., verhagen, j., wagenmakers, e.j., 2015. an introduction to bayesian hypothesis testing for management research. journal of management 41, 521–543. doi:https://doi.org/doi/10.1177/0149206314560412. barton, s.b., sanford, a.j., 1993. a case study of anomaly detection: shallow semantic processing and cohesion establishment. memory & cognition 21, 477–487. doi:https://doi.org/10.3758/bf03197179. bender, e.m., gebru, t., mcmillan-major, a., shmitchell, s., 2021. on the dangers of stochastic parrots: can language models be too big?, in: proceedings of the 2021 acm conference on fairness, accountability, and transparency, pp. 610–623. doi:https://doi.org/10.1145/3442188.3445922. van benthem, j., 2008. logic and reasoning: do the facts matter? studia logica 88, 67–84. proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 213 https://doi.org/10.3765/elm https://www.elm-conference.net/ doi:https://doi.org/10.1007/s11225-008-9101-1. brown, t.b., mann, b., ryder, n., subbiah, m., kaplan, j., dhariwal, p., neelakantan, a., shyam, p., sastry, g., askell, a., agarwal, s., herbert-voss, a., krueger, g., henighan, t., child, r., ramesh, a., ziegler, d.m., wu, j., winter, c., hesse, c., chen, m., sigler, e., litwin, m., gray, s., chess, b., clark, j., berner, c., mccandlish, s., radford, a., sutskever, i., amodei, d., 2020. language models are few-shot learners. url: https://arxiv.org/abs/ 2005.14165. bürkner, p.c., 2017. brms: an r package for bayesian multilevel models using stan. journal of statistical software 80, 1–28. calvillo, d.p., bratton, j., velazquez, v., smelter, t.j., crum, d., 2022. elaborative feedback and instruction improve cognitive reflection but do not transfer to related tasks. thinking & reasoning 0, 1–29. doi:https://doi.org/10.1080/13546783.2022.2075035. christiansen, m.h., chater, n., 2016. the now-or-never bottleneck: a fundamental constraint on language. behavioral and brain sciences 39, e62. doi:https://doi.org/10.1017/s0140525x1500031x. christianson, k., 2016. when language comprehension goes wrong for the right reasons: goodenough, underspecified, or shallow language processing. quarterly journal of experimental psychology 69, 817–828. doi:https://doi.org/10.1080/17470218.2015.1134603. christianson, k., luke, s.g., ferreira, f., 2010. effects of plausibility on structural priming. journal of experimental psychology: learning, memory, and cognition 36, 538–544. doi:https://doi.org/10.1037/a0018027. claidière, n., trouche, e., mercier, h., 2017. argumentation and the diffusion of counter-intuitive beliefs. journal of experimental psychology: general 146, 1052–1066. doi:https://doi.org/10.1037/xge0000323. clark, e., august, t., serrano, s., haduong, n., gururangan, s., smith, n.a., 2021. all that’s ‘human’ is not gold: evaluating human evaluation of generated text, in: proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing (volume 1: long papers), association for computational linguistics. pp. 7282–7296. doi:https://doi.org/10.18653/v1/2021.acllong.565. cook, a.e., walsh, e.k., bills, m.a., kircher, j.c., o’brien, e.j., 2018. validation of semantic illusions independent of anomaly detection: evidence from eye movements. quarterly journal of experimental psychology 71, 113–121. doi:https://doi.org/10.1080/17470218.2016.1264432. cook, p., stevenson, s., 2010. no sentence is too confusing to ignore, in: proceedings of the 2010 workshop on nlp and linguistics: finding the common ground, pp. 61–69. url: https://aclanthology.org/w10-2109. dale, r., 2021. gpt-3: what’s it good for? natural language engineering 27, 113–118. doi:https://doi.org/10.1017/s1351324920000601. dale, r., viethen, j., 2021. the automated writing assistance landscape in 2021. natural language engineering 27, 511–518. doi:https://doi.org/10.1017/s1351324921000164. dankers, v., lucas, c.g., titov, i., 2022. can transformer be too compositional? analysing idproceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 214 https://doi.org/10.3765/elm https://www.elm-conference.net/ iom processing in neural machine translation. url: https://arxiv.org/abs/2205. 15301. devlin, j., chang, m., lee, k., toutanova, k., 2018. bert: pre-training of deep bidirectional transformers for language understanding. corr abs/1810.04805. url: http://arxiv. org/abs/1810.04805, arxiv:1810.04805. elkins, k., chun, j., 2020. can gpt-3 pass a writer’s turing test? journal of cultural analytics 5, 17212. doi:https://doi.org/10.22148/001c.17212. erickson, t.d., mattson, m.e., 1981. from words to meaning: a semantic illusion. journal of verbal learning and verbal behavior 20, 540–551. doi:https://doi.org/10.1016/s00225371(81)90165-1. ferreira, f., 2003. the misinterpretation of noncanonical sentences. cognitive psychology 47, 164–203. doi:https://doi.org/10.1016/s0010-0285(03)00005-7. ferreira, f., patson, n.d., 2007. the ‘good enough’ approach to language comprehension. language and linguistics compass 1, 71–83. doi:https://doi.org/10.1111/j.1749818x.2007.00007.x. finlayson, m., mueller, a., gehrmann, s., shieber, s., linzen, t., belinkov, y., 2021. causal analysis of syntactic agreement mechanisms in neural language models. arxiv preprint arxiv:2106.06087 doi:https://doi.org/10.48550/arxiv.2106.06087. floridi, l., chiriatti, m., 2020. gpt-3: its nature, scope, limits, and consequences. minds and machines 30, 681–694. doi:https://doi.org/10.1007/s11023-020-09548-1. fortuin, e., 2014. deconstructing a verbal illusion: the ‘no x is too y to z’ construction and the rhetoric of negation. cognitive linguistics 25, 249–292. doi:https://doi.org/10.1515/cog2014-0014. frazier, l., clifton, jr, c., 2015. without his shirt off he saved the child from almost drowning: interpreting an uncertain input. language, cognition and neuroscience 30, 635–647. doi:https://doi.org/10.1080/23273798.2014.995109. frederick, s., 2005. cognitive reflection and decision making. journal of economic perspectives 19, 25–42. doi:https://doi.org/10.1257/089533005775196732. generative pretrained transformer, g., thunström, a.o., steingrimsson, s., 2022. can gpt3 write an academic paper on itself, with minimal human input? url: https://hal. archives-ouvertes.fr/hal-03701250. gibson, e., bergen, l., piantadosi, s.t., 2013. rational integration of noisy evidence and prior semantic expectations in sentence interpretation. proceedings of the national academy of sciences 110, 8051–8056. doi:https://doi.org/10.1073/pnas.121643811. helwe, c., clavel, c., suchanek, f.m., 2021. reasoning with transformer-based models: deep learning, but shallow reasoning, in: 3rd conference on automated knowledge base construction, p. https://openreview.net/forum?id=ozp1wrgtf5_. hossain, m.m., kovatchev, v., dutta, p., kao, t., wei, e., blanco, e., 2020. an analysis of natural language inference benchmarks through the lens of negation, in: proceedings of the 2020 conference on empirical methods in natural language processing (emnlp), association for computational linguistics, online. pp. 9106–9118. doi:10.18653/v1/2020.emnlp-main.732. hosseini, a., reddy, s., bahdanau, d., hjelm, r.d., sordoni, a., courville, a.c., 2021. proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 215 https://doi.org/10.3765/elm https://www.elm-conference.net/ understanding by understanding not: modeling negation in language models. corr abs/2105.03519. url: https://arxiv.org/abs/2105.03519. hupkes, d., dankers, v., mul, m., bruni, e., 2020. compositionality decomposed: how do neural networks generalise? journal of artificial intelligence research 67, 757–795. doi:10.48550/arxiv.1908.08351. kahardipraja, p., madureira, b., schlangen, d., 2021. towards incremental transformers: an empirical analysis of transformer models for incremental nlu. corr abs/2109.07364. url: https://arxiv.org/abs/2109.07364. kassner, n., schütze, h., 2019. negated and misprimed probes for pretrained language models: birds can talk, but cannot fly. arxiv preprint arxiv:1911.03343 doi:10.18653/v1/2020.aclmain.698. kejriwal, m., santos, h., mulvehill, a.m., mcguinness, d.l., 2022. designing a strong test for measuring true common-sense reasoning. nature machine intelligence 4, 318–322. doi:https://doi.org/10.1038/s42256-022-00478-4. lieber, o., sharir, o., lenz, b., shoham, y., 2021. jurassic-1: technical details and evaluation. white paper. ai21 labs. linzen, t., baroni, m., 2021. syntactic structure from deep learning. annual review of linguistics 7, 195–212. doi:https://doi.org/10.1146/annurev-linguistics-032020-051035. liu, y., ott, m., goyal, n., du, j., joshi, m., chen, d., levy, o., lewis, m., zettlemoyer, l., stoyanov, v., 2019. roberta: a robustly optimized bert pretraining approach. url: https://arxiv.org/abs/1907.11692. makowski, d., ben-shachar, m.s., lüdecke, d., 2019. bayestestr: describing effects and their uncertainty, existence and significance within the bayesian framework. journal of open source software 4, 1541. doi:https://doi.org/10.21105/joss.01541. marcus, g., davis, e., 2020. gpt-3, bloviator: openai’s language generator has no idea what it’s talking about. mit technology review url: https://www.technologyreview.com/2020/08/22/1007539/ gpt3-openai-language-generator-artificial-intelligence-ai-opinion. mccoy, r.t., pavlick, e., linzen, t., 2019. right for the wrong reasons: diagnosing syntactic heuristics in natural language inference. arxiv preprint arxiv:1902.01007 doi:https://doi.org/10.18653/v1/p19-1334. michel, p., neubig, g., 2018. mtnt: a testbed for machine translation of noisy text, in: proceedings of the 2018 conference on empirical methods in natural language processing, association for computational linguistics, brussels, belgium. pp. 543–553. url: https://aclanthology.org/d18-1050. o’connor, e., 2015. comparative illusions at the syntax-semantics interface. los angeles, ca: university of southern california dissertation . paape, d., 2021. the role of incremental and superficial processing in the depth charge illusion: experimental and modeling evidence. url: psyarxiv.com/jp2ma. psyarxiv preprint. paape, d., vasishth, s., von der malsburg, t., 2020. quadruplex negatio invertit? the on-line processing of depth charge sentences. journal of semantics 37, 509–555. doi:https://doi.org/10.1093/jos/ffaa009. proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 216 https://doi.org/10.3765/elm https://www.elm-conference.net/ passban, p., saladi, p.s., liu, q., 2020. revisiting robust neural machine translation: a transformer case study. arxiv preprint arxiv:2012.15710 doi:https://doi.org/10.48550/arxiv.2012.15710. patel, a., bhattamishra, s., goyal, n., 2021. are nlp models really able to solve simple math word problems? arxiv preprint arxiv:2103.07191 doi:https://doi.org/10.48550/arxiv.2103.07191. r core team, 2022. r: a language and environment for statistical computing. r foundation for statistical computing. vienna, austria. url: https://www.r-project.org/. radford, a., wu, j., child, r., luan, d., amodei, d., sutskever, i., 2019. language models are unsupervised multitask learners. openai blog 1, 9. reder, l.m., kusbit, g.w., 1991. locus of the moses illusion: imperfect encoding, retrieval, or match? journal of memory and language 30, 385–406. doi:https://doi.org/10.1016/0749596x(91)90013-a. rosenthal, j.s., 2008. monty hall, monty fall, monty crawl. math horizons 16, 5–7. url: https://www.jstor.org/stable/25678763. ryu, s.h., lewis, r.l., 2021. accounting for agreement phenomena in sentence comprehension with transformer language models: effects of similarity-based interference on surprisal and attention. arxiv preprint arxiv:2104.12874 doi:https://doi.org/10.48550/arxiv.2104.12874. sobieszek, a., price, t., 2022. playing games with ais: the limits of gpt-3 and similar large language models. minds and machines 32, 341–364. doi:https://doi.org/10.1007/s11023022-09602-0. stan development team, 2022. stan modeling language users guide and reference manual. url: https://mc-stan.org. version 2.26.8. thomson, k.s., oppenheimer, d.m., 2016. investigating an alternate form of the cognitive reflection test. judgment & decision making 11, 99–113. uchendu, a., ma, z., le, t., zhang, r., lee, d., 2021. turingbench: a benchmark environment for turing test in the age of neural text generation. arxiv preprint arxiv:2109.13296 doi:https://doi.org/10.48550/arxiv.2109.13296. vaswani, a., shazeer, n., parmar, n., uszkoreit, j., jones, l., gomez, a.n., kaiser, ł., polosukhin, i., 2017. attention is all you need, in: guyon, i., luxburg, u.v., bengio, s., wallach, h., fergus, r., vishwanathan, s., garnett, r. (eds.), advances in neural information processing systems, curran associates, inc. doi:https://doi.org/10.48550/arxiv.1706.03762. vos savant, m., 1997. the power of logical thinking. st. martin’s press, new york. wason, p.c., 1968. reasoning about a rule. quarterly journal of experimental psychology 20, 273–281. doi:https://doi.org/10.1080/14640746808400161. wason, p.c., reich, s.s., 1979. a verbal illusion. quarterly journal of experimental psychology 31, 591–597. doi:https://doi.org/10.1080/14640747908400750. wolf, t., debut, l., sanh, v., chaumond, j., delangue, c., moi, a., cistac, p., rault, t., louf, r., funtowicz, m., davison, j., shleifer, s., von platen, p., ma, c., jernite, y., plu, j., xu, c., le scao, t., gugger, s., drame, m., lhoest, q., rush, a., 2020. transformers: stateof-the-art natural language processing, in: proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, association for computational linguistics, online. pp. 38–45. url: https://aclanthology.org/2020. proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 217 https://doi.org/10.3765/elm https://www.elm-conference.net/ emnlp-demos.6. zhang, y., ryskin, r., gibson, e., 2022. a noisy-channel approach to depth-charge illusions. preprint available at ssrn: https://ssrn.com/abstract=4130042 . proceedings of elm 2: 202-218, 2023 dario paape: when transformer models are more compositional than humans. 218 https://doi.org/10.3765/elm https://www.elm-conference.net/ causal selection: the linguistic take elitzur a. bar-asher siegal, noa bassel and york hagmayer* abstract. causal selection is a widely discussed topic in philosophy and cognitive science, concerned with characterizing the choice of “the cause” among the many individually necessary and jointly sufficient conditions which any effect depends on. in this paper, we argue for an additional selection process underlying causal statements: causative construction selection (cc-selection), which pertains to the choice of linguistic constructions used to express causal relations. we aim to answer the following question: given that a speaker wishes to describe the relation between one of the conditions and the effect, which linguistic constructions are available? we take cc-selection to underlie causal selection, since the latter is restricted by the linguistic possibilities resulting from the former. based on a series of experiments, we demonstrate that factors taken previously as contributing to causal selection should, in fact, be considered as the parameters that license the various linguistic constructions under given circumstances, based on previous knowledge about the causal structure of the world (the causal model). these factors are therefore part of the meaning of the causative expressions. keywords. causation; causal selection; causative construction selection; lexical semantics; causal reasoning 1. introduction. the double selection problem. the occurrence of any event requires many different conditions to hold (mill 1884, a system of logic, volume i, chapter 5, §3). more precisely, many conditions are individually necessary and only jointly sufficient in order for a target event to take place, with several sets of jointly sufficient conditions relating to any event kind (mackie 1965). these conditions may include other events, states and constant background conditions, intentional actions or unintentional behaviors by an agent, and/or properties of the patient. to take a very simple example, the opening of an automatic door may depend on one sufficient set of conditions including electricity, the door being unlocked, and an agent pressing the door-open button. another sufficient set may include a door handle, the door being unlocked, and an agent pushing the handle. imagine a situation in which a person walks up to the door, pushes the button and the door opens. an observer, who wishes to describe what happened, is faced with what we call the double selection problem, involving causal selection on the one hand, and causative-construction selection on the other. the first problem has been widely discussed in philosophy and the cognitive sciences: the observer has to decide which among the many necessary and – in the particular situation – jointly sufficient conditions was the cause of the door opening. many theoretical accounts have been proposed in the recent years, alongside empirical studies testing how people select the cause from a set of conditions (cheng & novick 1991, hilton 1990 inter alia). studies show, for example, that an action by an agent that violates social norms is more likely to be considered the cause of an * authors: elitzur a. bar-asher siegal, the hebrew university of jerusalem (ebas@mail.huji.ac.il), noa bassel, the hebrew university of jerusalem (noa.bassel@mail.huji.com) and york hagmayer, university of göttingen (york.hagmayer@bio.uni-goettingen.de). research for this paper was supported by the state of niedersachsen, germany, "forschungskooperation niedersachsen-israel” for the project "talking about causation: linguistic and psychological perspectives" given to first and third authors with nora boneh. we thank nora boneh and anne temme for their assistance in early stages of the study. proceedings of elm 1: 027-038, 2021 c©2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer published by the lsa with permission of the author(s) under a cc by license. 27 https://doi.org/10.3765/elm https://www.elm-conference.net/ effect than an action which is within normative conventions, even in a situation where both actions are necessary for the effect to take place (hitchcock & knobe 2009, reuter et al. 2014, and icard, kominsky & knobe 2017 inter alia). the second problem has attracted far less attention in the cognitive sciences: in order to differentiate between causes, an observer has to describe the causal relations through an available linguistic construction, which involves what we call the causative construction selection (henceforth ccselection). in any causal statement, the speaker selects, along with the cause, a linguistic causative construction which appropriately describes the relation behind the observed course of events. in the above example, an observer may state: the pushing of the button opened the door, or pushing the button caused the door to open, to name just two possibilities. the array of causative constructions available includes connectives like because (of), from, by, overt causative verbs like make and cause, and change-of-state (cos) verbs such as open and boil, and other options available across languages (see bar-asher siegal & boneh 2020 for a definition of “causative constructions”. for typologies of causative constructions see shibatani 1976, comrie 1981 and song 1996). the cc-selection problem can be phrased in different ways. we will treat it as answering the following question: given that a speaker wishes to describe the relation between one of the conditions and the effect, which linguistic constructions are available? notably, this question does not assume singularity of causes. that is, with respect to each of the causative constructions, it is possible that more than one condition can be described as the cause, and our question aims at discovering the range of linguistic expressions that are available. cc-selection has largely been ignored in philosophy and psychology, although its relevance has been demonstrated by research inspired by linguistic theories (e.g. wolff 2003). a parallel question has been raised within theoretical linguistics (cf. dowty 1979), where various analyses correlate cos causatives like mary opened the door with direct rather than indirect causation. an example of the latter would be pushing somebody who accidently fell against the door open button (fodor 1970, shibatani 1976, wolff 2003 inter alia). further semantic differences between the various causative constructions have been recently raised by bar-asher siegal & boneh (2020) (for discussion on differences between specific construction see also neeleman & van der koot 2012, maienborn & herdtfelder 2017, bar-asher siegal & boneh 2019, nadathur & lauer 2020). we take cc-selection to be more crucial in the choice of a statement than causal selection, due to the fact that causal selection is restricted by the linguistic availabilities resulting from cc-selection. consider again the door example: determining whether “the cause” of the door to open was electricity, the person or the pushing of the button requires these possibilities to be stated. the relation between causal selection and cc-selection can be observed in experimental studies on causal selection (e.g. knobe and fraser 2008): first the participant is confronted with a causal scenario, whose underlying causal structure is known or provided, and in which (usually) two events occur, followed by a target event. the participants are then presented with causal statements, which generally include the phrases event a caused the target event and event b caused the target event, and asked to indicate how much they agree with the given statements. in this common experimental paradigm, the researcher pre-selects the causative constructions for the statements they regard as appropriate descriptions. hence the researcher has made an act of cc-selection, based upon which the participants are asked to make an act of causal selection. in this paper, we demonstrate that cc-selection affects causal judgements, showing first which causative constructions are available to observers describing the causal dependency between a condition (a member of a sufficient set of conditions) and an effect (section 2). we explore the semantics of causative verbs based on structural equation models, and propose a respective formal proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 28 https://doi.org/10.3765/elm https://www.elm-conference.net/ theory. second, we present a series of experiments (section 3), in which we investigate ccselection, showing that participants systematically rate the acceptability of certain causative constructions higher under specific conditions. we test how factors that have been shown to affect causal selection (cf. danks 2017) affect cc-selection, including violation of social norms and foreseeability of the effect by the involved agents. we show that these factors can be considered as the parameters that license the linguistic construction under given circumstances (section 4). these factors, therefore, become part of the meaning/truth-conditions of the causative expressions. 2. a theory of the semantics of causative constructions. in linguistics, the object of investigation into casual statements has traditionally been the structural and interpretative properties of causative constructions. nonetheless, dealing with causation is not trivial within formal approaches to semantics. the challenge has to do with the fact that formal approaches to the semantics of natural languages are truth-conditional and model-theoretic. in such frameworks, the meaning of a sentence is taken to be the proposition which is true or false relative to some model of the world. it is not trivial, however, to model causal statements, as they do not describe simple state-of-affairs in the world, or even in possible worlds. a handle: =1 if handle is turned; else =0 b lock: =1 if door is locked; else =0 c circuit: =1 if closed; else =0 d electricity: =1 if running; else =0 e button: =1 if pressed; else =0 f door opens: =1 if opens; else =0 g button =1 ⊨ circuit =1 sufficient set h (automatic opening): circuit =1 & electricity =1 & lock =0 ⊨ door opens =1 sufficient set i (manual opening): handle =1 & lock =0 ⊨ door opens =1 figure 1: structural equations and graphical models of two sufficient sets of conditions for an effect f. over the last decade, several works have explored the approach developed by judea pearl (pearl 2000) in the context of computer science to examine causality through structural equation modeling (sem), as a way to provide a model for the truth conditions of causal statements (baglini & francez 2016, baglini & bar-asher siegal 2020, nadathur and lauer 2020). in sem, causality is modeled by graphs that fit networks of constructs to data. on this approach dependencies between states of affairs are represented as a set of pairs of propositions and their truth values. here, we rely on baglini & bar-asher siegal’s (2020) formal definition for causal models. considering once more the example of the automatic door, we can define the variables (pairs of propositions and truth values) in a-f in figure 1. the fact that some variables depend on others for their value is represented by structural entailments in g-i. variables can be classified as belonging to one of two types: exogenous variables do not depend on any other variable (in the model). the values of the endogenous variables, in contrast, are based on the values of variables on which they depend. in our door example, the exogenous variables are a, b, d and e. the endogenous variables are c and f. dependencies within the sem can be represented qualitatively with directed acyclic graphs model (as in figure 1). nodes correspond to variables, and arrows indicate the direction of dependency: the value of an originating node dictates the value of nodes it points to. sufficient sets are circled (cf. vanderweele & robins (2009). h i proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 29 https://doi.org/10.3765/elm https://www.elm-conference.net/ following mackie (1965), we treat the nodes in the causal model as inus conditions: a variable or a set of variables which are insufficient but necessary alone, but together unnecessary but sufficient. in our example, when the door is unlocked, a closed circuit with electricity supplied constitutes a set which is sufficient but not necessary (since the pair of conditions unlocked and handle together are also sufficient for opening the door). we follow baglini & bar-asher siegal’s (2020) formal system for the definitions of necessary conditions, a sufficient set of conditions and situations within the sem framework. (a situation is defined as a set of pairs of propositions s in a language p and their values.) in this approach, the sem encodes speakers’ knowledge of the causal structure. moreover, formal definitions of various types of nodes/conditions (such as inus), can be used to capture the requirements for licensing linguistic judgments, or in our terms – for defining cc-selection. following baglini & bar-asher siegal (2020), we take a cos verb applied to a certain condition q representing the cause in the model (“the pushing of the button opened the door”),1 which is part of a situation s, to yield an acceptable causal statements under the conditions in (1). similarly, (2) captures the licensing conditions when q is the subject of an overt causative cause (“pushing the button caused the door the open”): (1) ∃q∃e∃t∃s:suff(s)m,r = 1 & (q ∈ s)m & s(e) & τ(e) ⊆ t & ∀t’ < t∀e’ : τ (e’) ⊆ t’ → [¬q(e’)] (2) ∃q∃e∃t∃s:suff(s)m,r = 1 & (q ∈ sm & q(e)) the function suff(icient) takes a situation (s) – a set of pairs of propositions and their values – and returns 1 if it is a sufficient set in the model for a specific result (r). the formula amounts to a description of a completion event. thus, in this line of analysis, lexical causatives select the temporally last condition to complete the sufficient set of conditions as its subject (1), while the periphrastic causative “cause” selects any condition in the set (2) (i.e., any inus condition). baglini & bar-asher siegal (2020) demonstrate that these formal descriptions capture previous observations regarding “direct causation” in the literature. the next section presents an investigation of the claims represented by (1)-(2), in a variety of experiments, showing that while it holds to a large degree, it must be slightly modified. 3. experiments 3.1 aim and hypotheses. we report on a series of 3 experiments aiming to empirically investigate cc-selection, by measuring the effect of the semantics of causative constructions on the acceptance of causal statements. based on the theory described in the previous section, we hypothesized that (1) statements with cos verbs will be more accepted for conditions completing a sufficient set of (preexisting) conditions; (2) sentences with an overt cause to construction will only be sensitive to whether their subject is an inus condition for the effect to take place. 3.2 overview. we confronted participants with a common effect structure in which two conditions conjunctively generated the target effect. hence, the two conditions were inus conditions in the terms of mackie (1965). participants were presented with various scenarios, in which two conditions were generated independently one after the other before the effect took place. we manipulated the temporal order such that both conditions were equally necessary for the effect, but only the condition occurring second completed a sufficient set. participants were asked to rate causal statements referring to each condition individually. one type of statement used a lexical, cos causative (e.g., suzan opened the window), the second type a periphrastic, overt causative (suzan caused the 1 in sem conditions are represented as propositions. we follow a long tradition since dowty (1979) according to which the dps in the actual causal statements are “representatives” of these propositions. proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 30 https://doi.org/10.3765/elm https://www.elm-conference.net/ window to open). experiment 1 was designed to show that observers prefer certain causative constructions to describe what happened (i.e. cc-selection). experiment 2 explored how the violation of social norms affects cc-selection. experiment 3 did the same with respect to foreseeability. 3.3. experiment 1. design. the study had a 2 (order of causes) x 2 (causative construction) x 4 (scenario) design. all three factors were manipulated within subjects. the study was run anonymously online using limesurvey (limesurvey.org). according to regulations at the university of goettingen, no clearance by an ethics committee was required. participants. we collected data of 35 participants, 32 of which passed the comprehension test. adult participants were recruited from prolific (prolific.org). english had to be their first language. materials and procedure. first, participants were informed that we are interested in how people use and understand language and that they would be presented with various scenarios and asked several questions examining their understanding of the scenario. then, participants were explicitly asked to indicate their informed consent to participate. next participants were presented with the first of four scenarios (expose/drawings, flood/land, open/door, or set off/alarm). (3) presents the the set off/alarm scenario. the participants were asked to rate four statements according to their compatibility with the facts presented in the scenario, as in (4). (3) the kagan family has a motion-sensitive security system, which they switch on when they leave the house. last monday, mary switched the system on, not knowing that her daughter, emily, was staying at home. when emily woke up, she passed in front of one of the motion sensors and activated it. the alarm went off. (4) (a) mary set off the alarm. (b) emily set off the alarm. (c) mary caused the alarm to go off. (d) emily caused the alarm to go off. the rating scale ranged from 1 (least compatible) to 7 (perfectly compatible). after their answer, participants were queried about the order of events to check whether they correctly grasped the given information. participants continued to the next scenario without receiving any feedback. the order of the scenarios and the order of the presented statements was randomized. statistical analysis. in all experiments we used a multilevel model to analyze the data, taking into account the design of the study. factors (here order, causative construction, scenario, and the interaction of order and causative construction) were entered as fixed effects. figure 2: results of experiment 1. mean ratings of causal statements and 95% confidence intervals are shown. higher ratings indicate higher acceptance. proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 31 https://doi.org/10.3765/elm https://www.elm-conference.net/ results. figure 2 displays the mean ratings depending on order and causative construction for the four scenarios. across all scenarios, ratings of statements with cos causatives were highly sensitive to order. ratings were much higher when the condition completed the sufficient set. ratings for statements involving overt causatives were less sensitive to order. statements with the first causal condition being the subject were rated higher when an overt causative construction was used than when a cos causative was. for the set off/alarm scenario this means that they rated the statement “mary set off the alarm” as less acceptable than “mary caused the alarm to go off”. there was no difference for statements referring to the second cause, which completed the sufficient set. this pattern was confirmed by the statistical analysis. the multilevel model allowed to predict participants’ ratings, loglik = –41.9, chi2 = 83.7, p < .0001, r2 = .15. the main effect contrast of order was significant, t(474) = 9.45, p <.0001, as was the main effect contrast of causative, t(474) = 3.32, p = .001, and their interaction, t(474) = 3.66, p =.0003. discussion. the findings provide evidence for cc-selection: participants rated the acceptance of causative constructions differently depending on which causal condition the statement referred to. when the causal condition completed the sufficient set, overt and cos causatives were considered appropriate. by contrast, cos causatives were considered less appropriate than overt causatives for a necessary condition that did not complete the sufficient set. these findings support our first hypothesis: participants were highly sensitive to the completion of a sufficient set when cos verbs were used. regarding the second hypothesis, we found that participants were less sensitive to a completion of a sufficient set for overt causatives. 3.4 experiment 2. in this experiment, we investigated the interaction between cc-selection and violation of social norms, which has been shown to strongly affect causal selection (cf. knobe & fraser 2008, icard et al. 2017). we hypothesized that a violation would have a stronger impact on the acceptance of statements when the subject represents the first condition with overt causatives than with cos causatives. design. the study had a 2 (order of causes) x 2 (causative construction) x 3 (scenario)2 x 2 (first agent violates norm vs. second agent violates norm) design. while the first three factors were manipulated within participants, the last factor (violation) was manipulated between participants. again, the study was run anonymously online. participants. seventy-four people participated (37 per violation condition). five participants were excluded, because they failed the comprehension test. recruitment and selection criteria were the same as in experiment 1. materials and procedure. instructions and the procedure were the same as in experiment 1. three new scenarios involving two agents were presented (lock/computer, set off/alarm, burst/ tank). the lock/computer scenario is presented in (5), with two possible continuations in (a-b). participants were asked to rate the statements in (6) on a scale from 1 (do not agree at all) to 7 (completely agree). after providing ratings, participants’ understanding of the scenarios was tested. the order of the scenarios and the order of the presented statements was randomized. (5) the cyber defense company iforce has a secured server which allows only one user to be logged into its system at a time. if a second user tries to log in, the system locks itself. according to schedule the senior developer beth works on the system every day between 7:00 and 13:00. her team-mate frank is scheduled to work on the same system from 13:15 until 19:00. 2 a fourth scenario did not involve agents violating norms and is, therefore, not reported. proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 32 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) first agent violates norms: …last week, beth didn’t pay attention to the time and stayed logged in past 13:00. frank logged in at his regular hour. the system locked. (b) second agent violates norms: … last week, beth worked her regular hours. frank felt that he was behind on his tasks and logged into the system at 12:58. the system locked. (6) (a) beth locked the system, (b) frank locked the system, (c) beth caused the system to lock. (d) frank caused the system to lock. results. the results are depicted in figure 3. when the second agent violated the norm (lower row in figure 3), the first cause received very low ratings. the second cause (completing the sufficient set) received high ratings regardless of the causative construction. by contrast, when the first agent violated the norm (upper row in figure 3), ratings were sensitive to order and causative construction.3 statements referring to the second cause were rated similarly in the respective scenario given both causative constructions. statements referring to the first cause, however, were accepted more when an overt causative was used. the latter findings replicate the findings of experiment 1. we also replicated the findings that overt causative statements referring to agents violating norms are rated higher than those referring to agents not violating norms (hitchcock & knobe 2009, reuter et al. 2014 and icard, kominsky & knobe 2017 inter alia). importantly, our findings show that there is an interaction of causative construction, order, and norm violation. figure 3: results of experiment 2. mean ratings of causal statements and 95% confidence intervals are shown. higher ratings indicate higher acceptance. the descriptive findings were corroborated by the statistical analysis. the multilevel model allowed to predict participants’ ratings, loglik = –228.9, chi2 = 457.9, p < .0001, r2 = .42. the following effects were significant: main effect of order (f (1,750) = 70.4, p < .0001), main effect of causative (f (1,750) = 10.6, p = .001), main effect of violation, (f (1,750) = 7.57, p = 0.006), interaction of order and causative (f (1,750) = 9.55, p = .002), interaction of order and violation (f (1,750) = 488.0, p < .0001), and the three-way interaction (f (1,750) = 18.9, p < .0001). 3 there is some inter-scenario variation, the analysis of which is beyond the scope of this paper. proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 33 https://doi.org/10.3765/elm https://www.elm-conference.net/ discussion. these findings again show the significance of cc-selection: participants rated different causative constructions differently with respect to the causal condition (here, an agent) the statement referred to. they also indicated a subtle interaction of norm violation, order, and causative construction. preferences for a specific construction was strongest when the statement referred to an agent acting first and violating a norm. in this case, an overt causative was clearly preferred over a cos causative. the fact that participants sometimes preferred a cos causative for the first cause (not completing the sufficient set) over a cos causative for the second effect (completing the sufficient set) indicates that our hypotheses need to be modified. a violation of norms affected cc-selection beyond the completion of a sufficient set. 3.5 experiment 3. in experiment 2, the norm violator could have foreseen the action of the other agent, but the agent conforming to the norm could not. foreseeability has been shown to moderate the effect of norm violation in causal selection (reuter et al. 2014). therefore, the aim of experiment 3 was to explore the effect of foreseeability on cc-selection. design. the study had a 2 (order of causes) x 2 (causative construction) x 4 (scenario) x 2 (first agent foresees action of second agent vs. first agent does not foresee action of second agent) mixed design. while the first three factors were manipulated within participants, foreseeability was manipulated between participants. participants. ninety-four people participated (47 per violation condition). three participants failed the comprehension test and were excluded from the analysis. recruitment and selection criteria were the same as in experiment 1. materials and procedure. instructions and procedure were the same as in the previous experiments. four scenarios were presented to participants (lock/computer, stop/elevator, set off/alarm, open/door). the lock/computer scenario is given in (7), with two possible continuations in (a-b). note that there was no explicit social norm that forbade the first agent to act. participants were asked to rate on a scale from 1 to 7 how much they agreed with the statements in (8). (7) the cyber defense company iforce has a secured server “f1”, which allows only one user to be logged into its operation system at the same time. if a second user tries to log in, the system locks itself. frank is the programmer responsible for performing daily checks on the f1 system, every day at 2pm. (a) first agent foresees second action: beth is frank’s old teammate and knows about his usual work schedule. last monday, beth logged into the system at 1:45pm, knowing that frank would log in later. frank logged in from his computer at his regular hour. the operation system locked. (b) first agent does not foresee second action: last monday, on her first day at work, frank’s new team-mate beth logged into the system at 1:45pm, not knowing that frank will log in later. frank logged in from his computer at his regular hour. the operation system locked. (8) (a) beth locked the system. (b) frank locked the system. (b) beth caused the system to lock. (d) frank caused the system to lock. results. results are displayed in figure 4. on the left-hand side, the results for the individual scenarios are shown, on the right-hand side the averages across scenarios. as in the previous two experiments, participants preferred statements with an overt causative over a statement with a cos causative when the statement referred to the first necessary but not sufficient cause. across scenarios, there was no clear preference for a particular causative construction for statements referring proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 34 https://doi.org/10.3765/elm https://www.elm-conference.net/ to the second cause. note that there was an effect of foreseeability for overt and for the cos constructions. the difference between ratings for the first and the second agent was smaller, when the first agent could foresee the action of the second agent (see figure 4 right hand side). figure 4: results of experiment 3 per scenario (left-hand side) and across scenarios (right-hand side). mean ratings of causal statements and 95% confidence intervals are shown. higher ratings indicate higher acceptance. the multilevel model allowed to predict participants’ ratings, loglik = –78.4, chi2 = 156.8, p < .0001, r2 = .10. the following effects were significant: main effect of order (f (1,1355) = 72.8, p < .0001), main effect of causative (f (1, 1355) = 36.1, p < .0001), main effect of foreseeability (f (1, 1355) = 5.89, p = 0.015), scenario (f (1, 1355) = 6.84, p = 0.001), interaction of order and causative (f (1, 1355) = 17.5, p < .0001), interaction of order and violation (f (1,750) = 488.0, p < .0001), and the interaction of order and foreseeability (f (1, 1355) = 16.6, p < .0001). followup analyses of the ratings for cos showed that there was a significant effect of foreseeability (f (1,631) = 4.10, p = .043) and an interaction of order and foreseeability (f (1,631) = 4.09, p = .043). the same analysis for overt causatives yielded no main effect of foreseeability (p=.18), but a strong interaction effect of order and foreseeability (f (1,631) = 14.2, p = .0002). discussion. again, we found cc-selection to be crucial. when the first agent was the subject of the sentence, participants preferred a statement with an overt causative over a lexical causative regardless of whether the first agent could foresee the action of the second agent. when the statement referred to the second agent, there was no clear preference for a particular causative construction. foreseeability affected the acceptance with respect to the first agent for overt causatives and cos verbs, the latter to a lower degree. there is an important limitation to the experiments: we manipulated the completion of a sufficient set through temporal order. therefore, one might argue that the results show that cos causatives are merely sensitive to order (see einhorn & hogarth 1986, henne et al. 2021for the effect of order on causal judgments). an ongoing trial tests this possibility. however, note that there is a theoretical motivation for why cos causatives should be sensitive to a completion of the sufficient set (baglini & bar-asher siegal 2020), and notably, sensitivity to the completion of a sufficient set entails sensitivity to temporal order. 4. the semantics of the overt causative “cause” and cos verbs. we can now consider properties taken previously as contributing factors to causal selection as parameters in the licensing of proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 35 https://doi.org/10.3765/elm https://www.elm-conference.net/ linguistic constructions under given circumstances. accordingly, these factors are taken as part of the meaning/truth conditions of the linguistic expressions. our results show that temporal order, and thereby the completion of a sufficient set, had an effect on both types of constructions, contra to our second hypothesis, and against the common claim relating direct causation with cos causatives and not with overt ones. in other words, while the findings are in line with the “direct causation” analysis of lexical causatives, the effect of temporal order on the overt causative is unexpected. norm violation and foreseeability showed interactions with construction and order, which means that these factors affect the acceptance of causative constructions differentially. table 1 summarizes the interactions of order, norm violation and foreseeability with linguistic construction. it reveals variation between scenarios (always/sometimes) and also relative influence between the two constuctions. the results show that speakers’ evaluations of the adequacy of different causal statements vis à vis a particular state of affairs vary systematically, depending on the type of linguistic expression employed to describe them. this variation indicates that we must treat ccselection independently, and that causal selection depends on linguistic facts (i.e. the choice of constructions) and not merely on the metaphysical or cognitive characteristics of the relata. change-of-state verbs relative influence overt cause to order (completion of a sufficient set) always a factor > always a factor violation of norms sometimes a factor < always a factor foreseeability always a factor < always a factor table 1: interactions of factors and linguistic constructions across experiments 1-3 following these results, we suggest to revise the proposal of baglini & bar-asher siegal (2020) reviewed in section 2, regarding the selection constraints for both types of constructions. while being an inus condition is probably the basic semantic requirement for using the construction with cause to, it is not enough. while it is possible that there are various factors that license the use of this construction, it is possible to offer one systematic principle, for the licensing of the cause to constructions. the higher sensitivity to norm-violation (experiment 2) and to the ability of agents to foresee the effect (experiment 3) both pertain to the degree of responsibility attributed to the condition with respect to the effect (see sytsma et al. 2012 and samland and waldmann 2016 for the notions of moral responsibility and blame in the context of causal selection). we propose that in assigning the role of the cause in the causative construction (i.e., the subject of the sentence), speakers seek to blame the specific condition for the occurrence of the effect. blame can be naturally assigned due to responsibility, but also as a result of a completion of a sufficient set. accordingly, an event is perceived as more “responsible”, or “blameworthy”, for an effect if it is the last to complete the set of sufficient conditions (see henne et al. 2021 regarding the notion of “recency”). if this proposal is on the right track, all factors are criteria for the same constraint: the condition represented by the subject of the cause-construction must be perceived as the one which is more responsible than the other according to at least one parameter. we therefore suggest (9) as a representation of an additional constraint to that in (2) on the choice of condition (q) among all conditions (cs): (9) ∀c∈s (c≠q responsibility (q) > responsibility (c)) with respect to the cos construction, we see that the requirement that the condition represented by the subject completes the sufficient set is stronger with this construction. however, we must account for two additional facts: in experiment 2, norm violation was a factor for accepting this proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 36 https://doi.org/10.3765/elm https://www.elm-conference.net/ construction, and in experiment 3 we found that, when the agent in the first condition could foresee the effect, this condition received a higher rating. in light of this, we wish to make the following preliminary proposal according to which foreseeability is also related to the notion of completion of the sufficient set. thus, there are two modes of completion of a sufficient set: an objective take: the last event which completes a sufficient set (then only time order matters); and a subjective take: the last condition that the agent didn’t know would be fulfilled (hence foreseeability matters). this difference is formally captured between (10), which repeats (1), and (11), in which the model is indexed with a perspective of a certain individual (i below). (10) ∃q∃e∃t∃s:suff(s)m,r = 1 & (q ∈ s)m & s(e) & τ(e) ⊆ t & ∀t’ < t∀e’ : τ (e’) ⊆ t’ → [¬q(e’)] (11) ∃q∃e∃t∃s:suff(s)m,i,r = 1 & (q ∈ s)m,i & s(e) & τ(e) ⊆ t & ∀t’ < t∀e’ : τ (e’) ⊆ t’ → [¬q(e’)] according to this, when the agent forsees that the second condition will take place, his own action is the last condition that he cannot know would be fulfilled. therefore, his action subjectively completes the sufficient set. in this way we can explain the results from experiment 3. in experiment 2, the agent violating the norm could expect the occurence of the other condition, therefore it might also be a case of a subjective completion of a sufficent set. 5. conclusion. three experiments demonstrated the significance of cc-selection, by showing that the acceptance of a causal statement was affected by the choice of the causative construction. consequently, we took the factors affecting the acceptance of the various constructions to be parameters that license the linguistic construction under given circumstances. we propose that they are components in the meaning of the causative expressions. an important ramification from these results is that future studies in cognitive science on causal selection must control for the linguistic construction used to express causative relations, thus accounting also for cc-selection. references baglini, rebekah, and elitzur a. bar-asher siegal. 2020. “direct causation: a new approach to an old question.” u. penn working papers in linguistics 26. baglini, rebekah, and itamar francez. 2016. “the implications of managing.” journal of semantics 33 (3): 541–60. https://doi.org/10.1093/jos/. bar-asher siegal, elitzur a., and nora boneh. 2019. “sufficient and necessary conditions for a non-unified analysis of causation.” in proceedings of the 36th west coast conference on formal linguistics, edited by richard stockwell, maura o’leary, zhongshi xu, and z.l. zhou, 55–60. http://www.lingref.com/cpp/wccfl/36/index.html. ———. 2020. “causation: from metaphysics to semantics and back.” in perspectives on causation: selected papers from the jerusalem 2017 workshop, edited by elitzur a. barasher siegal and nora boneh, 3–51. jerusalem studies in philosophy and history of science. cham: springer international publishing. https://doi.org/10.1007/978-3-030-343088_1. cheng, patricia w., and laura r. novick. 1991. “causes versus enabling conditions.” cognition 40: 83–120. comrie, bernard. 1981. language universals and linguistic typology. oxford: blackwell. danks, david. 2017. “singular causation.” the oxford handbook of causal reasoning, 201–15. dowty, david. 1979. word meaning and montague grammar. dordrecht: reidel. einhorn, j.hillel, and robin hogarth. 1986. “judging probable cause.” psychological bulletin 99: 3–19. proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 37 https://doi.org/10.3765/elm https://www.elm-conference.net/ fodor, jerry a. 1970. “three reasons for not deriving ‘kill’ from ‘cause to die.’” linguistic inquiry 1: 429–38. henne, paul,aleksandra kulesza, karla perez and augustana houcek. 2021. “counterfactual thinking and recency effects in causal judgment.” cognition 212. https://doi.org/10.1016/j.cognition.2021.104708 hilton, denis. 1990. “conversational processes and causal explanation.” psychological bulletin 107: 65–81. hitchcock, christopher, and joshua knobe. 2009. “cause and norm.” journal of philosophy 106: 587–612. icard, thomas f., jonathan f. kominsky, and joshua knobe. 2017. “normality and actual causal strength.” cognition 161 (april): 80–93. https://doi.org/10.1016/j.cognition.2017.01.010. knobe, joshua, and ben fraser. 2008. “causal judgment and moral judgment: two experiments.” moral psychology 2: 441–47. mackie, john l. 1965. “causes and conditions.” american philosophical quarterly 2: 245–64. maienborn, claudia, and johanna herdtfelder. 2017. “eventive vs. stative causation: the case of german causal von-modifiers.” linguistics and philosophy 40: 279–320. mill, john stuart. 1884. a system of logic, ratiocinative and inductive: being a connected view of the principles of evidence and the methods of scientific investigation. vol. 1. longmans, green, and company. nadathur, prerna, and sven lauer. 2020. “causal necessity, causal sufficiency, and the implications of causative verbs.” glossa: a journal of general linguistics 5 (1). neeleman, ad, and hans van koot. 2012. “the linguistic expression of causation.” in the theta system: argument structure at the interface, edited by martin marijana marelj everaert and tal siloni, 20–51. oxford: oxford university press. pearl, judea. 2000. causality: models, reasoning, and inference. cambridge, ma: cambridge university press. reuter, kevin, lara kirfel, raphael van riel, and luca barlassina. 2014. “the good, the bad, and the timely: how temporal order and moral judgment influence causal selection.” frontiers in psychology 5. https://doi.org/10.3389/fpsyg.2014.01336. samland, jana, and michael r waldmann. 2016. “how prescriptive norms influence causal inferences.” cognition 156: 164–76. shibatani, masayoshi. 1976. the grammar of causative constructions (syntax and semantics 6. edited by ed. new york: academic press. song, jae jung. 1996. causatives and causation: a universal-typological perspective. london: longman. sytsma, justin, jonathan livengood, and david rose. 2012. “two types of typicality: rethinking the role of statistical typicality in ordinary causal attributions.” studies in history and philosophy of science part c: studies in history and philosophy of biological and biomedical sciences 43 (4): 814–20. vanderweele, tyler j., and james m. robins. 2009. “minimal sufficient causation and directed acyclic graphs.” the annals of statistics 37 (3): 1437–65. wolff, phillip. 2003. “direct causation in the linguistic coding and individuation of causal events.” cognition 88: 1–48. proceedings of elm 1: 027-038, 2021 elitzur a. bar-asher siegal, noa bassel and york hagmayer: causal selection: the linguistic take. 38 https://doi.org/10.3765/elm https://www.elm-conference.net/ investigating a shared mechanism in the priming of manner and quantity implicature joe cowan & napoleon katsos* abstract. in the current paper, we investigate the existence of a shared derivation mechanism between manner and quantity implicature. as per the gricean-inspired perspective, both manner and quantity implicature are derived in a substantially analogous fashion, relying on the consideration of alternative ways in which the speaker could have spoken, but didn’t. in contrast, other accounts (e.g., grammatical accounts) of quantity implicature consider manner implicature and quantity implicature to be distinct in their derivational mechanisms. previous studies have found that quantity implicature can prime the derivation of subsequent quantity implicature both within and between quantity implicature subtypes in a structural priming paradigm, suggesting that ad hoc, numeral and some quantity implicature are governed by the same derivational mechanism. we have applied a structural priming paradigm to the case of manner implicature to investigate 1) whether manner implicature can be primed, 2) whether manner implicature can prime manner implicature and 3) whether manner implicature can be primed by quantity implicature. through manner-manner priming, the paper addresses the psycholinguistic reality of manner. while quantity-manner priming probes the existence of a shared derivational mechanism between the phenomena. we show that manner implicature can prime manner implicature under certain experimental circumstances and that ad hoc quantity, but not some quantity implicature can also prime manner implicature, whereas some quantity implicature cannot. keywords. manner implicature; quantity implicature; scalar implicature; priming. 1. introduction. debate exists concerning the mechanism that drives the derivation of quantity implicature. as per accounts based on grice’s (1975) framework, quantity implicature (qi) can be considered a pragmatic phenomenon and a constituent of the wider concept of conversational implicature. below are examples qi: (1) a: ‘some of the cars are green’ +> not all the cars are green. (2) a: did you meet john and bill? b: ‘i met john’ +> i did not meet bill. * authors: joe cowan, university of cambridge (jc2280@cam.ac.uk) & napoleon katsos, university of cambridge (nk248@cam.ac.uk). proceedings of elm 2: 36-48, 2023 c©2023 joe cowan and napoleon katsos published by the lsa with permission of the author(s) under a cc by license. 36 https://doi.org/10.3765/elm https://www.elm-conference.net/ as per a gricean program, the utterance ‘some of the cars are green’ is logically compatible with a reality in which all the referenced cars are green and a reality in which some, but not all, of the referenced cars are green. as such, there exists a resulting ambiguity, one that must be elucidated by pragmatic means. to disambiguate the utterance, a listener is said to assume that a more informative, relevant statement must be false, for, if it were true, a rational speaker would have surely used it. for instance, the phrase ‘all of the cars are green’ is more informative than the utterance ‘some of the cars are green’. if the latter were used to denote a reality in which all referenced cars were green, the speaker would be violating grice’s cooperative principle; the mutually assumed commitment to a set of conversational precepts, which, in part, enjoins a speaker to ensure optimal informativity of expression. therefore, upon hearing ‘some of the cars are green’ a listener is said to assume that the speaker means not[all of the cars are green] – a negation of the more informative alternative. then, not[all of the cars are green] is incorporated into the understanding of the articulated utterance ‘some of the cars are green’; the listener arriving at the understanding that the speaker means to communicate that ‘some but not all of the cars are green’. such is the derivation of quantity implicature as per gricean-inspired accounts. naturally, there are competing alternatives that posit the derivation of quantity implicature as a phenomenon occurring on a compositional level. chierchia (2006) and fox (2007) and chierchia, fox & spector (2012) provide accounts of qi that drastically depart from the gricean program, explaining quantity implicature to be a grammatical phenomenon, whereby a silent only operator is present in the syntax of scalars with an analogous meaning to the overtly expressed ‘only’. take the following: (3) ‘nancy took some of the toys.’ (4) ‘nancy took only some of the toys.’ as per a syntax-oriented explanation, the enriched ‘some but not all’ interpretation of some is represented by the covert insertion of a silent o (only) operator: (5) nancy took o[some of the toys] turning now to manner implicature (mi), per gricean-inspired accounts, manner implicature (mi) can be considered an analogous phenomenon; both qi and mi can be said to arise via violation of grice’s cooperative principle and reasoning about alternative utterances the speaker could have said. the maxim of manner, a constituent of grice’s cooperative principle, states that speakers must avoid obscurity and ambiguity of expression and ensure apt utterance brevity, therefore avoid prolix or obtuse utterances. take the following example: (6) a: ‘can you cook?’ b: ‘i am able to mix ingredients together to form something edible’ +> i cannot cook well. here, as with qi, a listener is said to understand that the speaker means not the alternative, in this case, ‘i can cook’. however, crucially, unlike qi, this alternative is not more informative, rather, less marked than the articulated utterance. levinson (2000) repackages grice’s conversational maxims related utterance form (ambiguity and brevity) to explain this via the m heuristic. the m heuristic comprises of two precepts: one that pertains to the speaker, the other to the listener. according to the m heuristic, the speaker adheres to a maxim whereby to reference an abnormal, unusual, or stereotypical situation, one must use marked language that contrasts with proceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 37 https://doi.org/10.3765/elm https://www.elm-conference.net/ the unmarked language with which one would typically reference a normal corresponding situation. simultaneously, a listener is to adhere to the acknowledgment that marked expressions pertain to marked situations (levinson, 2000). in sum, marked language = marked situation. under the gricean-inspired perspective, qi and mi are derived via the same mechanism; one that relies on consideration, and subsequent negation of, alternative ways in which a speaker could have spoken, but chose not to. qi and mi however diverge regarding the nature of the alternative. in qi, the alternative is more informative while in mi it is less marked. moreover, whereas in qi the meaning of the alternative is negated, in mi the meaning of the alternative is not negated by the mi, rather the implicated meaning, i.e., the stereotypical connotations of the utterance. therefore, when choosing to use the prolix ‘i am able to mix ingredients together to form something edible’ a speaker could intend to negate the stereotypical connotations of ‘i can cook’: ‘i can cook edible food’, ‘i can cook visually appealing food’ or ‘i can cook food considered typical or normal of my culture’, etc. despite the similarities between the nature of the alternative, as per a gricean-inspired perspective substantial similarities exist between the derivation of qi and mi. in contrast, the mechanisms governing qi and mi derivation diverge fully if we are to consider qi a grammatical phenomenon. given that mi concerns not what is said, rather how it is said, mi is necessarily post-compositional and cannot be explained by the existence of a grammatical operator. qi, on the other hand, is derived by a silent o, a grammatical operator dedicated to scalar implicature. taking stock, the relation between qi and mi is up for debate; either the derivation of qi and mi implicature are governed by the same pragmatic mechanism, or qi and mi are governed by distinct mechanisms. moreover, given that a compositional account of qi needn’t preclude any subsequent pragmatic enrichment, it could also be the case that qi and mi diverge in their origin – qi perhaps occurring on a syntactic level – but converge in their appeal to pragmatics to expound their meaning. 2. background. the nature of the mechanism governing the derivation of qi, and the degree to which this mechanism is shared between subtypes of qi, has been the subject of experimental exploration using structural priming techniques. bott & chemla (2016) investigate the extent to which priming can evidence a shared derivational mechanism between qi subtypes; some, numeral, ad hoc: (7) a: ‘some of the cats are sleeping’ +> not all the cats are sleeping. (some) (8) a: ‘i broke two of my fingers’ +> i broke exactly two of my fingers. (numeral) (9) a: ‘there is a dog in the garden’ +> there is a dog that isn’t mine in the garden. (ad hoc) according to bott & chemla, the above are all examples of qi. bott & chemla employed a reimagination of huang, spelke & snedeker’s (2013) ‘hidden box’ paradigm to investigate a theorized shared mechanism. the paradigm consisted of priming trials and critical trials in a prime #1 – prime #2 – critical trial order. in all trials, the participants were presented with a caption containing a qi, e.g., ‘there are some arrows’ and two images and were tasked with matching the caption to one of the displayed images. the images present in the trial were configurations of shapes that either did or didn’t correspond to the qi caption, e.g., a card with 6 arrows and 3 diamonds on it, or a plain card that said ‘better picture?’. proceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 38 https://doi.org/10.3765/elm https://www.elm-conference.net/ in the priming trials, participants were presented with weak primes or strong primes. the weak primes consisted of a caption containing a qi implicature, e.g., ‘there is a triangle’ (+> there is only a triangle) and two images, one image corresponding to the unenriched meaning, e.g., two triangles, and a false image that did not correspond to the caption, e.g., an image of a circle. for lack of a better option, in the weak prime trials, the participants were forced to select an image that matched not with the qi’s enriched meaning, rather the unenriched, non-qi meaning. in contrast, in the strong prime trials, the participants were presented 1) with an image that matched the unenriched qi, e.g., a triangle and a star, and 2) an image that matched the enriched qi, e.g., a single triangle. thus, in the strong prime trials, the participants were given free choice between either deriving or suspending the qi. given the salience of qi, the participants were exceedingly likely to select the image corresponding to the enriched qi. in the critical trials, the participants were faced with an image matching the unenriched qi and a blank image that read ‘better picture?’ – the implication being that the unseen referent of the ‘better picture’ card corresponds to the enriched qi. assuming the effect of structural priming, bott & chemla hypothesize that, in the strong prime condition, the participant’s selection of the image corresponding to the enriched qi, in the priming trial, is to prime the selection of the ‘better picture?’ card (the image tacitly corresponding to the enriched qi) in the succeeding critical trials. therefore, bott & chemla predict an increase in the selection of the ‘better picture?’ card in the critical trials in the strong prime condition. bott & chemla investigate a potential priming effect in two supra-conditions, firstly, within-category priming, whereby the same subtype of qi, e.g., some, is used throughout the prime #1 – prime #2 – critical trial sequence and a between-priming condition, whereby priming occurs between qi subtypes, e.g., a some prime #1 –some prime #2 – ad hoc critical trial sequence. the between-subtype condition investigates the extent to which there exists a shared derivational mechanism between qi subtypes – the rationale being that if one qi subtype primes another, there exists a shared mechanism. bott & chemla’s investigation indeed reports evidence to suggest that qi can prime qi of the same subtype, suggesting that prior derivation of one qi subtype can prime the subsequent derivation of the same qi subtype. bott & chemla also report this affect to hold between qi subtypes some, numeral, ad hoc, suggesting that there exists a shared derivational mechanism between qi subtypes. bott & chemla conclude that observed priming effect is likely resultative of the priming of an o-operator but is compatible with both gricean-inspired and grammatical accounts of qi. rees & bott (2018) present an adaptation of bott & chemla’s (2016) ‘better picture?’ paradigm with a priming condition that increases the salience of the qi implicature’s alternative (e.g., ‘some of the letters are b’’s alternative = ‘all of the letters are b’). in the alternative condition, participants were presented with a sentence containing a qi’s alternative (e.g., ‘all of the letters are b’) and two images, one corresponding to the sentence (e.g., an image with six letter bs) and one not corresponding to the sentence, but corresponding to the configuration of the enriched qi – the qi to which the sentence is the alternative (e.g., three letter bs and six letter as). crucially, the alternative prime does not require the explicit derivation of a qi. these alternative primes were tested against the weak and strong primes as per bott & chemla (2016) – the experiment investigating whether the use of qi itself, or the tacit salience of the qis alternative, is responsible for the demonstrated priming effect in bott & chemla (2016). rees & bott found that the use of an ‘alternative’ prime is just as effective as eliciting a priming effect as a strong prime in the proceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 39 https://doi.org/10.3765/elm https://www.elm-conference.net/ case of ad hoc, some and numeral qi, suggesting that the derivation of qi is not responsible for the observed efficacy of the strong prime in bott & chemla (2016) and rees & bott (2018). despite the important conclusions of bott & chemla (2016) and rees & bott (2018), a critical flaw exists in both studies: no baseline rate of qi implicature judgement is provided in either. to address this concern, marty, cowan, romoli, sudo and breheny (2021) presented a rerun of the priming experiments in bott & chemla (2016) and rees & bott (2018), providing novel baselines with which to assess the purported within-category priming effect. marty et al. find that while the strong prime is indeed a driver of a positive priming effect in the case of ad hoc primes when compared to the established baseline, there is no such increase in some qi, concluding that the purported increase of some qi after strong some primes is rather an effect of reverse priming on the part of the weak some prime. taken together, important conclusions can be drawn from the three studies: 1) both within and between-category priming exists for all three types of qi. 2) it is most likely the salience of the alternative driving the observed priming effect 3) the findings are compatible with a broad range of accounts of implicature (gricean-inspired, grammatical, a.o.) 3. current study. given that there exists theoretical debate concerning the extent to which qi and mi are analogous phenomena, it follows that structural priming paradigms may also prove worthwhile means with which to investigate mi and the relationship between qi and mi. as such, the current paper investigates two research questions. firstly: can mi be primed? – a question, which is intrinsically motivated, given that, to the extent of our knowledge, the structural priming of mi is novel. secondly: can qi cross-prime mi? – a question motivated by both its novelty and its ability to illuminate the existence of a shared derivational mechanism governing the derivation of qi and mi and to subsequently inform wider theoretical understanding of the phenomena as either distinct, approximate, or homologous. on a broader level, our endeavors contribute to the emergent body of work concerning the experimental investigation of mi. 4. methodology. the novel operationalization of mi in a structural priming paradigm represents a challenge; a researcher must create a set of functionally identical trials that reflect a necessarily ad hoc, context dependent phenomenon – a non-issue in the case of qi, whereby the manifestations of each reiteration of the qi follow an identical structure (e.g., some of the x are y). in a reimagination of the trials presented in bott & chemla (2016), rees & bott (2018) and marty, cowan, romoli, sudo and breheny (2021), the trials in the present study comprised of a caption involving a mi and two images. the type of mi used between the trials was one of markedness, whereby a marked, prolix description selects are marked, unusual referent. in terms of caption, the mi trials comprised of protracted definitions of quotidian, arbitrarily selected, objects, whereby mention of the common label of the object is circumvented (e.g., for the object ‘church’ the caption ‘select the picture with a building used for christian worship’ is used, see fig.1). the prolix definitions were abridged, and sometimes paraphrased, variations of google’s english dictionary entries, as provided by oxford languages (collected in march-may 2021). the entries were paraphrased to avoid the use of inappropriately formal language and were abridged so that they did not extend beyond a single sentence. in each trial, the caption was paired with two images: one representing a marked referent and the other an unmarked referent; the marked referent representing the enriched version of the mi, the marked-to-marked association. proceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 40 https://doi.org/10.3765/elm https://www.elm-conference.net/ the images were selected from google images under the discretion of the researcher, with [object label] and [weird object label] used as search terms. once a bank of unmarked – marked images were compiled, the items were subject to experimental pre-screenings. using online survey platform qualtrics, 45 native speakers of english were asked to provide a description of the 76 images, half were marked items and half were unmarked items in a between-participants presentation mode, ensuring that a participant would not see both the marked and unmarked variations of one object. one image appeared per trial, alongside the phrase ‘what do you see in this image?’ and a blank text box in which the participant could submit an answer. the purpose of running the item pre-screening was to ensure that the images selected were indeed appropriately associated with their label, ascertaining each images label rate allowed us exclude images with a < 70% label mention rate. the rationale for excluding such items that failed this criterion was that, to derive a mi, a participant must be aware that both images pertain to the same label, otherwise the markedness contrast – the crux of the mi – cannot be established. as such, the marked item needs to a be a solid referent of its preconceived label; expressing its markedness in a manner which does not compromise its concept belongingness (e.g., in selecting a cat unambiguous in its cat-ness, but nevertheless marked). resultingly, post item screening, we retained 18 sets of viable unmarkedmarked item pairings for use in the subsequent experiments. 5. baseline. to establish a point of reference from which to interpret the results of the proceeding priming studies, we measured the baseline rates of manner implicature derivation in the 18 selected unmarked-marked pairings. as explained, the experimental items consisted of a caption, a prolix description of the preconceived item label, and two images, both corresponding to the item label, with one of the images a typical example of the item label and the other an atypical example. the participants were asked to select which image best matched the caption with selection of the atypical item taken as indicative of the participant’s derivation of mi, i.e., an atypical caption to atypical image matching. all 18 item pairings appeared in a randomized, withinparticipants block that also included 10 filler items, whereby each participant provided a judgement on all 18 item pairings. 5.1 participants. the participant pool comprised of 63 adult monolingual english speakers recruited online via prolific.co. in all the experiments comprising the current study, participants who had previously participated in any other experiments involving the study were disallowed from participating. 5.2. results. fig.2 demonstrates the disparity in mi derivation rates. at the highest point of the spread, ‘hat’ has a baseline mi rate of 29.03%, while at the low end, ‘hot air balloon’ shows a figure 1: item 'church' proceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 41 https://doi.org/10.3765/elm https://www.elm-conference.net/ 0.00% 5.00% 10.00% 15.00% 20.00% 25.00% 30.00% 35.00% selection rate of atypical image over typical image across items figure 3: rate of mi derivation across items baseline mi deviation rate of 3.23%, with an average derivation rate of 11.83% (sd= 8.57%) across all experimental items. the items as they appeared to the experimental participants are shown in fig.3. 5.3. discussion. mi is, by nature, a heterogenous phenomenon and therefore as the context that guides mi varies, too does mi itself. as such, the baseline rate of mi derivation (taken as the selection rate of the atypical image over the typical image) varies drastically between item pairings. this is expected, as each different mi is essentially a one-off contrast-based implicature influenced by what is said, and how, and what is seen, and is incomparable in uniformity to qi, whereby the implicature relies on more uniform syntax and structure, e.g., ‘some of the x are y’ +> ‘not all xs are y’ in any imagination of the qi. the rate of mi derivation is not only heterogeneous but also low compared to the almost ceiling level rates of qi presented in marty et al. (2021). however, this it to be expected and is consistent with the notion that more inferential and context-dependent reasoning is require in activating the alternative expression in mi leading to reduced derivation (in some items) figure 2: items 'hat' and 'hot air balloon' proceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 42 https://doi.org/10.3765/elm https://www.elm-conference.net/ 6. experiment 1a. having established the baseline rate of mi derivation for the 18 prescreened item pairs, we conducted a within-category (mi > mi) priming experiment, addressing the question of mi’s ability to be primed. 6.1. structure. experiment 1a followed a filler – filler – prime #1 – prime #2 – critical trial order. the critical trials were identical to those presented in the baseline experiment; they consisted of a free choice between an atypical image and a typical image and a related caption. in contrast, the prime trials, while identical in caption and atypical image to the critical trials, differed in that the typical image was replaced by an unrelated, but equally typical, image of an object e.g., a typical car. this meant that in the prime trials, participants were confronted with a prolix caption and were forced to match it to a marked image, see fig.4 for an illustration, but were aware of the possibility or existence of unmarked objects within the experiment. figure 4: prime 'camera' the rationale governing this configuration was that the forced selection of the atypical image, and the subsequent derivation of mi, would prime the derivation of mi in the free-choice scenario in the succeeding critical trials à la bott & chemla (2016), rees & bott (2018) and marty et al. (2021). in experiment 1a, all item pairings were presented both as primes (with an adapted typical image) and as critical trials (as per baseline) between three trial blocks with participants pooled to one of the three trial blocks. the trial blocks were formulated in a manner whereby each participant ultimately provided judgement on 6 of the overall 18 critical trial pairings while the remaining 12 item parings functioned as primes. 6.2. participants. 180 adult monolingual english speakers recruited from prolific.co. 6.3. results. in the critical trials of experiment 1a, the rate of atypical image selection, taken as indicative of mi derivation, was numerically higher at 16.23% (sd = 12.34%) than our baseline at 11.86% (sd= 8.57%), but statistically insignificant as per a mann-whitney test u=126, p=0.261. 6.4. discussion. the derivation of mi in the prime trials has not significantly increased mi derivation in the critical trials. while it could be the case that mi cannot be primed, considering factors beyond this, the configuration of the prime trials in experiment 1a could be responsible for the observed lack of priming. there seems to be two reasons the prime trials may be deficient in experiment 1a, 1) with no certitude can it be said that the participants are engaging in mi in the prime trials; a participant could engage in semantic matching between the caption and the image, irrespective of whether the image is considered marked, circumventing the pragmatic engagement with the caption requisite in triggering mi derivation 2) the matching atypical improceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 43 https://doi.org/10.3765/elm https://www.elm-conference.net/ age may not be considered atypical, or marked, when not appearing next to the less marked, more typical alternative – it seems unlikely that the typicality of, say, a car, can consistently trigger the awareness of the atypicality of, say, a camera. 7. experiment 1b. considering the issues with the prime trials of experiment 1a, we conducted a reiteration with adapted priming trials. in experiment 1b, the priming trials were almost identical to the critical trials, except a highlighted caption reading ‘this image!’ was placed above the typical image, see fig.5. figure 5: updated item: 'camera' the experimental instructions were updated so that the participants were told ‘on some of the images you will be told which image the speaker means. please select the image they are talking about (it will be labelled 'this image!')’. the rationale for the adapted prime was that, by including the typical alternative, we contextualize the atypicality of the atypical image, rendering its markedness salient and that the forced selection of the atypical image avoids any purely semantic matching by the participant and instead requires the participant to appeal to pragmatics to rationalize the selection of the atypical image and ultimately arrive at the mi 7.1. participants. 180 monolingual english speakers recruited from prolific.ac.uk. 7.2. results. in the critical trials of experiment 1b, the rate of atypical image selection was 16.76% (sd = 8.81%), higher than our baseline of 11.86% (sd= 8.57%); a mann-whitney test indicating that this finding is statistically significant, u= 96, p= 0.036. 7.3. discussion. with the updated prime trial structure, we can observe a priming effect of the prime trials on mi derivation in the critical trials. while the increase is not dramatic, this could be explained by the pervasive elusivity of mi derivation. as discussed, mi derivation requires the concurrent consideration of various contextual factors, given that this is the case, it could be that, broadly, as derivation of mi is elicited more infrequently than that of qi (as evidenced by the baseline) even in a primed condition mi derivation remains infrequent. alternatively, it could be the case that, given the overall low rates of mi derivation, the operationalization of mi in the experimental trials is not robust enough to routinely trigger mi, both in baseline and primed conditions. given the unnaturalistic context of the trials and the lack of fleshed-out speaker identity, it would be surprising to see lower rates of mi than would be observed ‘in the wild’. moreover, in either case, it can be said that the increase in atypical image selection is indicative of an increase in mi derivation and, as such, we can conclude that 1) mi can be primed and 2) mi can be primed by mi. considering evidence of within-category priming of mi, to adproceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 44 https://doi.org/10.3765/elm https://www.elm-conference.net/ dress the questions concerning the relationship between mi and qi, a between-category experimental naturally follows. 8. experiments 2a & 2b. considering evidence of within-category priming of mi, to address the questions concerning the relationship between mi and qi, we conducted two betweencategory priming experiments using some and ad hoc qi to prime mi. 8.1. structure. as in experiments 1a & 1b, the trials followed a filler – filler – prime #1 – prime #2 – critical trial order, in which the critical trials were identical to those throughout the study at large. changed were the primes, presented in fig.5, whereby experiment 2a concerned some qi and 2b ad hoc implicature. in experiment 2a & 2b, the primes were identical in essence to those developed between bott & chemla (2016), rees & bott (2018) and marty et al. (2021); some primes consisted of two images and a caption. the caption took the form ‘select the picture in which some of the shapes are [shape]’. of the images, one matched the enriched qi, e.g., for ‘star’ an image of a card containing three stars and six squares, and the other matching the unenriched alternative e.g., all stars. see fig. 6. figure 6: 'ad hoc' and 'some' prime for experiment 2b, the ad hoc primes had the same caption-image configuration but instead read ‘select the picture with a [shape]’, with one image matching the enriched qi, i.e., a card containing one instance of a single shape, and the other the unenriched qi, i.e., a card containing two shapes, one pertaining to the caption and the other to a different shape. given the high baseline rates of qi derivation, the free choice between the enriched or unenriched qi presented on the cards self-selects for the image pertaining to the enriched interpretation and therefore acts as a prime without having to push participants to the selection (e.g., as in 1b with the ‘this image!’ caption). here, given that mi can be primed, if there is a shared derivational mechanism between qi and mi, we predict that qi derivation in the preceding prime trials will confer greater mi derivation in the succeeding critical trials. 8.2. participants. 360 adult monolingual english speakers recruited from prolific.co., split equally between studies. 8.3. results. in the ad hoc prime condition, the mi derivation rate increased to 18.04% (sd = 12.05%), a significant increase compared to a baseline of (11.86%) as per a mann-whitney test, u=93, p=0.028, in contrast, while we observed an increase from the baseline to 15.67% (sd = 9.95%) is observed in some qi prime condition, this is not found to reach statistical significance, u=110, p= 0.102. proceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 45 https://doi.org/10.3765/elm https://www.elm-conference.net/ 8.4. discussion. while we gathered evidence of cross-category priming exists, this can only be said to be the case with ad hoc qi and mi, not some qi and mi. the lack of a priming effect of some trials on mi could have been resultative of a more integral issue with the experimental structure, for instance, the lack of a robust conversational context in both the prime and mi critical trials when paired with the conventionalization of some qi, and a lack of contextual boundness not found in ad hoc qi, could have disengaged the participants from more profound consideration of speaker intention and thus failed to elicit the salience of mi in the eventual critical trials. while this could be the case, the difference between the ability of some and ad hoc qi to prime mi could be due to a more integral difference between the two types of qi and their relation to mi. if we are to consider the nature of the alternative in some and ad hoc qi and mi, it could be the case that the disparate levels of contextual consideration requisite in the search for the alternative are responsible for the disparate potency of some and ad hoc qi as primes. for instance, the derivation of some qi relies on the search for conventional alternatives, i.e., ‘most’ or ‘all’ which are then then negated. in contrast, as is naturally the case with mi, the alternative in the case of ad hoc qi is less conventionalized. for instance, take the following sentence and its potential ad hoc implicatures: (10) a: ‘i took bernie out for a walk’ +> i didn’t go to the shops/i didn’t walk rover. here, to construct the ad hoc scale with which to interpret the utterance, context is needed to determine whether the scale is a scale of tasks completed, or dogs walked. in (10) the alternative is not uniform between the potential ad hoc implicatures; ‘i took bernie out for a walk and did the shopping’ and ‘i took bernie and rover out for a walk’ being two alternatives. as such, the nature of ad hoc qi’s alternative approximates that of mi; both are reliant on context for their generation. therefore, it could be the case that the contextual boundness of ad hoc qi is what allows for the priming relationship between ad hoc qi and mi to occur and the lack of requisite contextual awareness in the generation of the alternative is where a priming effect between some qi and mi is thwarted. 9. overall discussion & conclusion. in general, in both baseline and prime conditions, we see a far lower rate of mi than can be observed for ad hoc and manner qi. as explained, this aligns with the understanding of mi as a context-dependent, one-off phenomenon; given that mi is derived from context and is not lexically or structurally triggered, it follows that in order for mi to be derived, a robust context, i.e., one in which mi derivation is motivated, is needed. despite this, in exp 1b, we did find evidence that mi derivation rates can be augmented by the existence of a mi prime in the context of our experimental setting. the existence of a withincategory mi priming effect gives credence to the psycholinguistic reality of markedness contrasts and to mi generally. in terms of cross-category priming, we observe a small cross-priming effect between ad hoc qi and mi implicature but not between some qi and mi, exp. 2a and 2b, see table 1 and fig,8, which is compatible with our understanding of the importance of the role of context in the derivation of these three types of implicature. the demonstration of a cross-priming effect is novel across pragmatic maxims, having previously been demonstrated between sub-types of qi (bott & chemla, 2016; rees & bott, 2018; marty, et al., 2021). while the existence of this crosspriming effect is compatible with both gricean-inspired and grammatical accounts of qi, the existence of a priming effect between ad hoc qi and mi suggests that ad hoc qi is, at least in proceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 46 https://doi.org/10.3765/elm https://www.elm-conference.net/ part, a pragmatic phenomenon and that a theoretical dichotomy between ad hoc qi and mi as truly distinct phenomena is unjustified. while the novelties of the presented study have provided evidence suggesting that mi can be primed by mi and by ad hoc qi, something that has theoretical ramifications, the same novelties have limited the presented study. given the rudimentary nature of the operationalization of mi in the trials, it is likely that the availability of mi derivation is being suppressed and that, with preceding iterations of priming studies involving mi, mi can be experimentally constructed in a way in which augments the baseline rate of mi, the effectiveness of mi as a prime and the susceptibility of mi as a subject of priming. a further question raised by the current study is the relationship between some qi and mi; is it the case that some qi cannot prime mi? or is it the case that, given the lesser degree of similarity between some qi and mi when compared to ad hoc qi and mi, that the experimental structure is inhibiting a potential some qi to mi priming effect? table 1: collated results figure 7: collated results 0.00% 5.00% 10.00% 15.00% 20.00% 25.00% 30.00% 35.00% baseline 1a: manner 1b: manner 2a: some 2b: ad hoc rate of mi across experiments experiment: rate of mi derivation significance in difference from baseline baseline: mi 11.83% n/a experiment 1a: mi – mi prime 16.23% p=0.261 experiment 1b: mi – mi prime 16.76% p= 0.036 experiment 2a: some qi – mi prime 15.67% p= 0.102 experiment 2a: ad hoc qi – mi prime 18.04% p=0.028 proceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 47 https://doi.org/10.3765/elm https://www.elm-conference.net/ bibliography: bott, l., & chemla, e. 2016. shared and distinct mechanisms in deriving linguistic enrichment. journal of memory and language, 91, 117–140. https://doi.org/10.1016/j.jml.2016.04.00 chierchia, g. 2004. scalar implicatures, polarity phenomena, and the syntax/pragmatics interface. in structures and beyond (pp. 39–103). oxford, united kingdom: oxford university press. chierchia, g., fox, d., & spector, b. 2012. scalar implicature as a grammatical phenomenon. in p. porter, c. maienborn, & k. von heusinger (eds.), an international handbook of natural language meaning (vol. 3, pp. 2297–2331). berlin, germany: de gruyter grice, h. p. 1975. logic and conversation. syntax and semantics 3: speech acts, 26-40 huang, y. t., spelke, e., & snedeker, j. 2013. what exactly do numbers mean? language learning and development, 9(2), 105–129. https://doi.org/10.1080/15475441.2012.658731 levinson, s. c. 1983. pragmatics (cambridge textbooks in linguistics). cambridge, united kingdom: cambridge university press levinson, s. c. 2000. presumptive meanings: the theory of generalized conversational implicature. the mit press. marty, p., cowan, j., breheny, r., romoli, j., & sudo, y. 2021. what primes what – an experimental framework to explore alternatives for sis. paper presentation presented at the 34th annual cuny conference on human sentence processing, new york , united states of america. proceedings of elm 2: 36-48, 2023 joe cowan and napoleon katsos: investigating a shared mechanism in the priming of manner and quantity implicature. 48 https://doi.org/10.3765/elm https://www.elm-conference.net/ transparency in the processing of temporal ambiguity: the case of embedded tense giuliano armenante, vera hohaus & britta stolterfoht* abstract. we report the results of one acceptability rating study and two self-paced reading studies on the form-meaning mismatch in the interpretation of past-under-past in complement clauses in english. across the three experiments, we find an off-line and on-line preference for the backward-shifted interpretation, in line with predictions of the structural approach to the ambiguity when assuming a processing preference for morphological transparent interpretation. keywords. tense semantics; consecutio temporum; sequence-of-tense parameter; form-meaning mismatches; ambiguity; processing strategies; comprehension 1. introduction. in english, embedded tenses in certain configurations give rise to ambiguities, most prominently in the case of past tense in a stative complement clause embedded under a pastmarked verb of reported speech. this past-under-past configuration is illustrated in (1). the sentence is compatible with two different direct utterances made by oliver, indicated in (1-a) and (1-b), which correspond to a backward-shifted reading (under which amber’s illness pre-dates oliver’s utterance) and a simultaneous reading (under which her illness temporally overlaps with his utterance). generalising broadly, structural approaches to this ambiguity treat sim as derived from back by additional morpho-syntactic technology that allows the lower past tense not to be interpreted as such (prominently, ogihara 1989, 1996, stowell 1996, kusumoto 1999, 2005). (1) oliver said that amber was sick. a. oliver said: “amber was sick.” (backward-shifted reading, back) b. oliver said: “amber is sick.” (simultaneous reading, sim) we sketch one possible implementation of such a view in (3), where the lower past-operator that we see in the logical form for back is deleted to derive sim, assuming that the use of past-tense morphology on the embedded verb can also be licensed by the past-operator in the matrix clause. in the absence of such an operator in the embedded clause and in interaction with the semantics of say in (2), amber’s illness and oliver’s utterance are interpreted to share an evaluation time. (2) for any possible world w ∈ dw, time t ∈ di, individual x ∈ de and tensed proposition p ∈ d⟨s,⟨i,t⟩⟩, j say (simplified) k(w)(t)(p) = 1 if and only if in all worlds w′ compatible with x’s utterance at t in w, p(w′)(t) = 1. *this research was funded by the deutsche forschungsgemeinschaft dfg (project b1, tübingen collaborative research centre 833, id # 75650358). the research underwent ethical review by the tübingen ethics committee (reference # 2017 0328 53) and the proportionate university ethics committee at the university of manchester (reference #s 2018-5463-7466, 2019-5463-10129). we would like to thank the audience of the 2019 tübingen “processing tense” workshop as well as petra augurzky, núria barrios-jurado, sigrid beck, maximilian berthold, chen-an chang, lilian gonzalez rodriguez, robin hörnig, zahra kolagar, alina mclellan, anne mucha, paul stott, rolf ulrich, and malte zimmermann. authors: giuliano armenate, universität potsdam (armenante@uni-potsdam.de), vera hohaus, the university of manchester (vera.hohaus@manchester.ac.uk) & britta stolterfoht, eberhard karls universität tübingen (britta.stolterfoht@uni-tuebingen.de). proceedings of elm 2: 1-12, 2023 c©2023 giuliano armenante, vera hohaus and britta stolterfoht published by the lsa with permission of the author(s) under a cc by license. 1 https://doi.org/10.3765/elm https://www.elm-conference.net/ (3) a. logical form for back: [ pastt∗ [ λt oliver sayw@,t [ λw λt′ pastt′ [ λt′′ amber sickw,t′′ ] ] ] ], where t∗ is the utterance time and w@ the actual world b. logical form for sim: [ pastt∗ [ λt oliver sayw@,t [ λw ————[ λt′′ amber sickw,t′′ ] ] ] ] in this paper, we investigate the processing predictions of this analysis. more generally, understanding how the processing system handles this particular ambiguity can also inform our understanding of the strategies the processing system adopts when there is no one-to-one mapping between form and meaning and multiple interpretations are available. 2. previous research. while considered a “touchstone for the adequacy of the semantics for tense” (von stechow 2009; 3), embedded tenses have received relatively little attention in the processing literature. while we find the experimental results in dickey (2000, 2001)’s landmark study to be overall inconclusive, they have been taken to point towards a processing preference for sim. the data from one of the adult-control groups in hollebrandse (2000)’s research on first language acquisition also suggests a slight acceptability preference for sim. gennari (2004) employs a design that relies on additional manipulations, but also observes an advantage for overlapping temporal intervals in reading times. mucha et al. (2022) in their comparison of english and polish, however, observe higher acceptability for back compared to sim for both languages. 3. research questions and experimental hypotheses. when it comes to processing, the structural approach outlined in the introduction can plausibly be taken to predict a processing preference for back over sim, given that an additional operation derives the former from the latter to allow the past tense morphology in the embedded clause not to be interpreted. more abstractly, this apparent mismatch between form and meaning is plausibly dispreferred. we translate these considerations here into the experimental hypothesis h1 in (4). (4) hypothesis h1 “what you see is what you get” (wysiwyg): comprehension is driven by morphological transparency. a past tense should initially always be interpreted as such, favouring back over sim. structurally, however, the resulting logical form underlying sim can also be argued to be simpler than the structure that derives back, which may result in a processing preference for sim. we summarise such a view in (5), as our alternative experimental hypothesis h2. (5) hypothesis h2 “representational simplicity”: comprehension is driven by structural simplicity at logical form. a past tense in the relevant configuration should initially not be interpreted, favouring sim over back. in the next section, we report two sets of experiments designed to test these two hypotheses using reading times and acceptability judgments. in general, we expect a processing preference to be reflected in shorter reading times and higher acceptability ratings. 4. experimental data. 4.1. experiments 1 and 2. the first two experiments investigated the two hypotheses for the type of configuration from the introduction, the temporal ambiguous interpretation of a clausal proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 2 https://doi.org/10.3765/elm https://www.elm-conference.net/ complement of a verb of saying with a past-marked, stative predicate. the two experiments share design and materials, with experiment 1 a reading-time study and experiment 2 a rating study, where a preceding context disambiguated between readings. design. both experiments adopted a 2x2 factorial design, with factors evaluation time (past vs future) and intended interpretation (back vs sim). evaluation time was manipulated within the visually presented context; see figure 1 for an example. the factor interpretation relates to the target items presented. a sample item from the two experiments is in (6), where the vertical bar indicates the presentation segments for experiment 1. figure 1: sample visual contexts for experiments 1 and 2 (6) a. context: oliver, yesterday: “amber was sick!” past,back target: oliver | said | that | amber | was | sick. b. context: oliver, yesterday: “amber is sick!” past,sim target: oliver | said | that | amber | was | sick. c. context: oliver, tomorrow: “amber was sick!” future,back target: oliver | will say | that | amber | was | sick. d. context: oliver, tomorrow: “amber is sick!” future,#sim target: oliver | will say | that | amber | was | sick. of particular interest for our research questions are the two past conditions in (6-a) and (6-b), for which an ambiguous past-under-past sentence is evaluated in a context that establishes back or sim. note however also that simultaneous readings are not available when embedding a pastmarked stative predicate under a future-marked verb of saying. as a result (and as indicated by # above), there is a mismatch between context and target for condition (6-d), which was designed to act as a control condition. predictions. of the hypotheses formulated earlier, h1 “wysiwyg” predicts a processing advantage for back, while h2 “representational simplicity” predicts an advantage for sim. for the first experiment, we expect an effect to arise on the auxiliary or possibly at the spillover region (= the adjective), since it is at this point that the processor can establish the temporal relation between the embedded and the matrix clause and align it with the preceding context. (we assume that the the temporal information retrievable from the context is used for ambiguity resolution; see also, for instance, trueswell & tanenhaus 1991.) under h1, we expect longer reading times for the non-transparent past,sim compared to past,back. h2 predicts that the reading times in the critical region a shorter for the structurally simpler past,sim, compared to past,back. for the second experiment, h1 predicts higher acceptability ratings for past,back compared to past,sim, while we expect past,sim to be rated better than past,back under h2. proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 3 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4.1.1. methods. materials. we constructed 48 sets of experimental items like (6), embedded among 40 fillers, for a total of 88 trials, with each participant being exposed to 12 experimental items per condition. each trial comprised a context picture paired with a target sentence. in order to minimize variation between items, all target sentences followed the rigid template in (7).1 (7) name1 said/will say that name2 was adjective. fillers contained reported speech constructions similar to the experimental items but targeting a different kind of aspectual ambiguity. we also designed 24 comprehension question relating to the names of the characters involved or the time of the story. participants. participants (n experiment 1 = 43, n experiment 2 = 40) were all native speakers of english. five participants from the initial 48 were excluded from experiment 1 because their accuracy in answering comprehension questions was below the 75%-threshold. participants received a cash reimbursement for their participation. procedure. experiment 1 was a moving-window self-paced reading study, while experiment 2 was an acceptability rating study. across experiments, participants first saw the context picture, followed by the target sentence. this was read word-by-word in the first experiment, while in the second experiment, participants were asked to rate the sentence based on its fit with the context picture on a scale from 1 “bad fit” to 6 “good fit”. of the trials, 3/11 ended with a comprehension question. each experimental session started with written instructions, followed by two practice trials containing relative clauses, and then two blocks of the 88 trials with a break halfway through. out of the 88 trials, four lists were formed following a latin square design, with stimuli presented in a randomised fashion. an average session took approximately 30 minutes, with participants tested individually at the university of manchester psycholinguistics laboratory. 4.1.2. results and discussion. the data analysis was conducted in the spss (statistical package for the social sciences) programming environment, performing a repeated-measures analysis of variance (anova). in our exploratory analyses, we speak of a significant effect only if the probability of making a type i error (α error) is below 5%. reading times. reading times were corrected for outliers first, computing only values within the 90-1,300 ms region. in a second step, reading times with 2.5 standard deviations away from the mean were excluded. the resulting mean values across all sentence regions are in figure 2. mean reading times for the two regions of interest, the embedded auxiliary and the adjective that followed it, are listed in table 1. the results show an effect of evaluation time for both regions of interest. at the auxiliary (that is, the embedded was), future attitude reports took significantly longer to read compared to past ones (f1,42 = 19.5, p < .001, η2 = .318). the same effect was found at the adjective (f1,42 = 9.6, p < .005, η2 = .186). by contrast, an effect of interpretation only reached significance at the adjective, with sim evoking longer reading times than back (f1,42 = 4.4, p < .05, η2 = .096), with no significant interaction between the two factors. zooming in on the 1if name1 was female, then name2 name was male, with the number of female and male names kept equal. only negative polar adjectives like furious, late, and drunk featured in the stimuli in all three experiments, as they were deemed to not bias against one reading over the other in the relevant configuration. proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 4 https://doi.org/10.3765/elm https://www.elm-conference.net/ two past conditions, we performed a t-test for evaluation time = past. while no difference was found at the auxiliary (t(42) = 1.187, p = .242), the effect approached significance at the adjective (t(42) = 1.939, p = .059), with back associated with faster reading times than sim. figure 2: mean reading times in ms for experiment 1, with say*= said/will say region condition mean reading time in ms standard deviation in ms embedded auxiliary past,back past,sim future,back future,sim 334.4 340.9 354.6 365.9 95.4 105.9 99.1 128.0 adjective past,back past,sim future,back future,sim 410.5 430.1 440.9 479.8 111.1 126.7 112.8 200.8 table 1: mean reading times for experiment 1 with standard deviations for regions of interest proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 5 https://doi.org/10.3765/elm https://www.elm-conference.net/ acceptability ratings. mean acceptability scores are listed in table 2.2 past future back 5.98 5.54 sim 5.38 4.16 table 2: mean values of participants’ ratings for experiment 2 (based on a scale from 1 “bad fit” to 6 “good fit”) overall, the future,sim control condition in experiment 2 received the lowest (but still relatively high) ratings, which show that the processor was sensitive to the temporal mismatch in this condition (despite the lack of a significant interaction in experiment 1). by contrast, past-under-past was rated the highest when the intended reading was backward-shifted (in the past,back condition), but the other interpretation (in the past,sim condition) still averaged a high rating. the statistical analysis revealed an effect of evaluation time (f1,47 = 244.9, p < .001, η2 = .839), with future conditions rated significantly worse than past. furthermore, we observe an effect of interpretation, in that participants rated back conditions significantly higher than sim conditions (f1,47 = 107.1, p < .001, η2 = .695). the interaction between evaluation time and interpretation is significant (f1,47 = 65.7, p < .001, η2 = .583), with back rated significantly higher than sim for future compared to past evaluation times. upon closer inspection, the difference between the two interpretations was also significant in the past conditions (t(39) = 4.921, p < .001), pointing to a preference for the backward-shifted reading. discussion. we cautiously take these results to indicate a processing preference for the morphologically transparent, backward-shifted interpretation of past embedded under past in complement clauses. the results thus lend initial support not only to an analysis of embedded tenses that derives the simultaneous reading from the backward-shifted reading, but also in favour of a wysiwyg processing strategy, as outlined in h1. one potential challenge for this conclusion relates to a potential priming effect, however: for past,back and future,back conditions, both the visual context as well as the target item contained the same form of the auxiliary to be (that is, was), while sim conditions involved the same lemma but different lexemes (that is, is in the visual context and was in the stimulus). we cannot exclude the possibility that this property of the design may have facilitated the processing of back conditions over sim conditions. the design of the experiment we present in the next section sought to address this concern. 4.2. experiment 3. we further tested h1 “wysiwyg” and h2 “representational simplicity” in a third experiment, in a design where the embedded past was locally ambiguous and the intended interpretation was not already retrievable from the preceding context. the design was intended to establish whether any processing preferences exists when it comes to back and sim and to address the methodological concerns raised in the discussion of experiments 1 and 2. it was also intended to provide insight into some aspects of how the ambiguity affects the incremental composition of temporal meaning in these cases. 2due to a coding error, five experimental items had to be removed from the analysis. proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 6 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4.2.1. design. the experiment tested ambiguous and non-ambiguous past complement clauses with different continuations, which served as disambiguation in the case of ambiguous past. the experiment adopted a 2x2 factorial design, with factors ambiguity (+amb vs -amb) and event (event1 vs event2), where the first event temporally preceded the second, and the second event coincided with the time at which the matrix verb was interpreted. we will explain the rationale behind this design in more detail for the sample item in (8), where the rehearsal and the concert make reference to the two events in question. (8) a. context: after last week’s final rehearsal, last night, +amb,event1 john’s band finally gave a concert, where i spoke to him about mary. . . (= back) target: john | said | that | mary | was sick, | so | that’s why | she | missed | the rehearsal | with | great | regret. b. context: after last week’s final rehearsal, last night, +amb,event2 john’s band finally gave a concert, where i spoke to him about mary. . . (= sim) target: john | said | that | mary | was sick, | so | that’s why | she | missed | the concert | with | great | regret. c. context: the other day john’s band taped their final rehearsal. -amb,event1 and when this morning i spoke to john about mary. . . target: john | thinks | that | mary | was sick | and | that’s why | she | missed | the rehearsal | with | great | regret. d. context: the other day john’s band finally gave a concert. -amb,event2 and when this morning i spoke to john about mary. . . target: john | thinks | that | mary | was sick | and | that’s why | she | missed | the concert | with | great | regret. in the +amb conditions, the context was set up in such a way as to introduce both the concert and the rehearsal that preceded it, and established that john and the speaker had a conversation at the concert. as a result, the target sentence without the that’s why-continuation in (8-a) and (8-b) is ambiguous, with the embedded clause plausibly allowing for both back and sim. this ambiguity is resolved by the rehearsal or the concert in the continuation, with the former only compatible with back and the latter only with sim. such an ambiguity does not arise in the -amb conditions, where the matrix verb is in the present tense; past-under-present embeddings do not allow for multiple temporal interpretations. the -amb conditions served as baseline conditions that allowed us to explore the time course of the ambiguity resolution and to control for effects of length and frequency associated with the nouns used for disambiguation. predictions. the design of the experiment was intended to allow us to explore the preferred temporal interpretation of ambiguous past (as reflected in reading times), but also the time course of ambiguity resolution, that is, when the processing system decided on the temporal interpretation of the locally ambiguous complement clause. does the processor delay committing to one of the interpretations as long as possible, in an attempt to avoid a potential revision later down the line? or is there a default preference for one reading that will allow for an interpretation to be determined soon as possible? from these questions, there derive three regions of interest for this experiment, namely, the embedded predicate (that is, was sick in (8) above), the disambiguating region (the proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 7 https://doi.org/10.3765/elm https://www.elm-conference.net/ rehearsal or the concert, for instance), and a spillover region, which was defined as the disambiguating region +1. under h1 “wysiwyg”, we expect an interaction at the disambiguating or the spillover region, with event2 evoking longer reading times than event1 in the +amb conditions but not in the -amb conditions. h2 “representational simplicity” predicts shorter reading times for the event2,+amb condition (= sim), compared to the event1,+amb condition (= back). in relation to the incrementality of the composition of the temporal meaning in these sentences, any such difference could either be a result of the processing system revising a previously assigned interpretation or the result of a dispreferred temporal interpretation that has been delayed but is now forced by the disambiguation. as for the embedded predicate, if the ambiguity goes unnoticed and the processing system commits to a temporal interpretation right away (likely, the preferred interpretation), we may not expect an effect of ambiguity. if both interpretations that are available at this point undergo more active consideration, we expect an increase in reading times for +amb conditions over -amb conditions. 4.2.2. methods. materials. 16 experimental items like (8) were constructed alongside 48 fillers, to yield a total of 64 trials. trials included a context sentence followed by a target sentence. of the trials, 3/8 ended with a comprehension question. disambiguated nouns were carefully paired so that they would depict events with a stereotypical temporal order (as in the case of a rehearsal and a concert), to facilitate the retrieval of the intended temporal interpretation. as in the first two experiments, the target sentences followed a fairly rigid template, with said in all the +amb conditions and thinks in all the -amb conditions. filler items exhibited a temporal or an aspectual local ambiguity, with a disambiguating continuation segment that was either a consecutive sentence similar to the one used for the experimental items or a relative clause. participants. 78 participants, native speakers of english, were recruited via the platform prolific. ten participants were excluded as they scored below 70% in the comprehension task, bringing the total number of participants to 68. participants were reimbursed for their participation. procedure. the experiment adopted a self-paced reading task, with contexts and targets displayed entirely masked on the screen together (see (8) above for the segmentation adopted), but separated by line breaks. three practice trials preceded the experimental session, which was split into four blocks. participants had the opportunity to take a break at the end of each trial. four lists were created following a latin square design, with stimuli presented in a randomised fashion. an experimental session took approximately 25 minutes to complete online. 4.2.3. results and discussion. reading times. results were analysed as in experiment 2.3 reading times were corrected for outliers, trimming values below 1,200 ms for the context and above 1,500 ms for the disambiguating region in the continuation. reading times were then log-transformed, with absolute values below a standard deviation of 3 from the mean being removed. the reading times computed for the two regions of interest to the temporal interpretation are reported in table 3. no effect of ambiguity or event was found at the disambiguating region. although not reaching significance 3due to a programming error, two items had to be removed from the analysis. proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 8 https://doi.org/10.3765/elm https://www.elm-conference.net/ (f1,67 = 3.127, p = .082, η2 = .045), we found a trend towards an interaction, with faster reading times for event1 (= the backward-shifted interpretation) than event2 (= the simultaneous interpretation) in the +amb condition compared to the -amb baseline; see also figure 3. no effect was observed at the spillover region (f1,67 = .729, p = .396, η2 = .011). lastly, there was no effect of ambiguity at the embedded predicate region (for the example in (8), was sick). there is thus no evidence that locally ambiguous sentences slowed down the processing system at the first point where it would have been possible to assign a temporal interpretation to the sentence and where the temporal ambiguity would have been detectable. region ambiguity event mean disambiguation +amb -amb event1 (= back) event2 (= sim) event1 event2 5.878 5.928 5.922 5.869 spillover +amb -amb event1 (= back) event2 (= sim) event1 event2 5.991 5.990 5.973 5.929 table 3: log-transformed mean reading times for regions of interests figure 3: reading times of participants for experiment 3 at the disambiguating region proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 9 https://doi.org/10.3765/elm https://www.elm-conference.net/ discussion. overall, these findings are compatible with the results from the previous two experiments and with a general processing preference for backward-shifted interpretations over simultaneous ones, in line with h1 “wysiwyg”. in relation to the on-line composition of the temporal interpretation, in the absence of an effect at the embedded predicate, it appears that the ambiguity did not slow down the processing system. it is plausible that the preferred interpretation was assigned as soon as possible and later revised if incompatible with the continuation. before we turn to the overall discussion, let us briefly comment on the lack of a significant interaction between ambiguity and event in this experiment, especially given the trend observed in the data. one possible explanation for this may be design-related, the other may be related to more general cognitive principles: in relation to the design, the sample size was likely not large enough to produce sufficient statistical power, given the small effect size (η2 = .45) and the challenge of constructing a higher number of good stimuli with the adopted design. from a cognition perspective, it has been observed that temporal remoteness is harder to process than temporal proximity, in that events that are temporally distant come with a processing cost compared to those closer to the speech time (see, for instance, gennari 2004). back involves an event that is temporally further away from both the utterance time to which the sentence is anchored and the time at which the matrix predicate is interpreted than the event involved in sim (that is, the rehearsal precedes the concert precedes the utterance time in the example item). since the events features prominently in the items, we may want to speculate that this cognitive bias was a confounding effect and curbed the magnitude of the expected interaction. a second, related confound might also have played an inhibiting role: more recently mentioned referents are easier to recover, and thus to process. within the design of experiment 3, these were by choice always noun phrases that would force sim. we opted against presenting the two events in a non-chronological order in the context for half of the trials to preserve naturalness and readability of the stimuli. 5. overall discussion. taken together, the three studies presented here provide initial evidence for an onand off-line preference for back over sim in the interpretation of ambiguous past-underpast. taken together, we cautiously take these results to provide support for h1 “wysiwyg” and a processing strategy that favours morphological transparency over structural simplicity to resolve form-meaning mismatches. these results challenge the often implicitly assumed preference for sim, but are in line with more recent findings in mucha et al. (2022). based on our findings in experiment 3, we have tentatively suggested that back may constitute a default interpretative strategy during on-line comprehension and that the ambiguous past in complement clauses does not cause the processing system to delay the composition of the temporal meaning of the sentence. these are preliminary findings, however, which should be investigated further in an adequately controlled experimental setting. while a preference for back is not compatible with h2 “representational simplicity” as discussed in section 3, the structurally simpler logical form for sim may not necessarily translate to a processing advantage over back if we discuss the two structural representations as a type of filler-gap constellation, with the embedded tense operator a filler and the temporal variable associated with the embedded predicate the gap. processing difficulty in filler-gap constellations is often measured in terms of the linear or structural distance between the filler and the gap, with proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 10 https://doi.org/10.3765/elm https://www.elm-conference.net/ increased distance in a dependency increasing processing load.4 assuming for the logical form underlying sim a reduced structure as in (3), which lacks a tense operator (that is, filler) in the embedded clause, the temporal variable (that is, the gap) associated with the embedded predicate will be bound by the matrix predicate (and in turn the matrix past), leading to a non-local chain. by contrast, the logical form underlying back is characterised by a local chain, which in return should translate into a processing advantage for back over sim. given these considerations, no preference would be expected for sim, even under a processing hypothesis based on the structural representations underlying the different readings. 6. conclusion. this paper investigated the processing of a prominent form-meaning mismatch in english that relates to the temporal interpretation of sentence. more specifically, we investigated the processing of past-tensed complement clauses embedded under a past-marked verb of saying/ attitude verb, which allow for both simultaneous and backward-shifted interpretations. overall, the results provide preliminary evidence in favour of a preference for back in both reading and rating tasks that prevents a processing delay in composing temporal meaning in these cases. we took these findings to provide empirical support for a processing strategy that hinges on a transparent mapping of morphological form to meaning where, by wysiwyg, past morphology by default translates to past meaning. in the light of these results, we reasoned that competing hypotheses built on structural simplicity may have to be revised, as the reduced temporal representations underlying sim come at the cost of a potential processing difficulty associated with their non-local filler-gap dependencies. beyond these findings, we hope that this paper will raise awareness for what we consider an under-researched area in semantic processing and highlight some of the methodological questions in researching the processing of embedded tense. directions for further research not only include addressing some of these methodological challenges with different designs, but also broadening the empirical footing of the research programme to include the processing of progressive-marked past predicates, which also give rise to back and sim (see, for instance, kusumoto 1999), and a wider array of embedded environments, including relative clauses and adjunct clauses. relative clauses in particular are compatible with a wider range of temporal configurations than complement clauses, without exhibiting structural ambiguity (see, for instance, kusumoto 2005, hohaus 2019). processing data from a wider range of languages would also be welcome in future research, especially given the range of well-documented cross-linguistic variation in the interpretation of embedded tenses (see, for instance, grønn & von stechow 2010, bochnak et al. 2019). references bochnak, m. ryan, vera hohaus & anne mucha. 2019. variation in tense and aspect, and the temporal interpretation of complement clauses. journal of semantics 36(3). 407–452. https://doi.org/10.1093/jos/ffz008. collins, chris. 1994. economy of derivation and the generalised proper binding condition. linguistic inquiry 25(1). 45–61. http://www.jstor.org/stable/4178848. 4for relevant discussion and references, see collins (1994), hamilton (1995), o’grady (1997), hawkins (1999, 2004), and o’grady et al. (2003), among many others. proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 11 https://doi.org/10.3765/elm https://www.elm-conference.net/ dickey, michael walsh. 2000. the processing of tense. amherst: university of massachusetts dissertation. dickey, michael walsh. 2001. the processing of tense: psycholinguistic studies on the interpretation of tense and temporal relations. dordrecht: kluwer. gennari, silvia p. 2004. temporal references and temporal relations in sentence comprehension. journal of experimental psychology: learning, memory, and cognition 30(4). 877–890. https://doi.org/10.1037/0278-7393.30.4.877. grønn, atle & arnim von stechow. 2010. complement tense in contrast: the sot parameter in russian and english. in atle grønn & irena marijanović (eds.), russian in contrast: grammar, 109–153. oslo: oslo studies in language 2(1). hamilton, robert l. 1995. the noun-phrase accessibility hierarchy in sla. in fred r. eckman, diane highland, peter w. lee, jean mileham & rita rutkowski weber (eds.), second language acquisition: theory and pedagogy, 101–114. mahwah: erlbaum. hawkins, john a. 1999. processing complexity and filler-gap dependencies across grammars. language 75(2). 244–285. https://doi.org/10.2307/417261. hawkins, john a. 2004. efficiency and complexity in grammars. oxford: oxford university press. hohaus, vera. 2019. the temporal interpretation of complement and relative clauses: contrasting english and samoan. journal of the southeast asian linguistics society 12(3). 42–60. hollebrandse, bartjan. 2000. the acquisition of sequence of tense. amherst: university of massachusetts dissertation. kusumoto, kiyomi. 1999. tense in embedded context. amherst: university of massachusetts dissertation. kusumoto, kiyomi. 2005. on the quantification over times in natural language. natural language semantics 13(4). 317–357. https://doi.org/10.1007/s11050-005-4537-6. mucha, anne, agata renans & jacopo romoli. 2022. sequence of tense and cessation implicatures: evidence from polish. natural language and linguistic theory online first. https://doi.org/10.1007/s11049-022-09545-2. ogihara, toshiyuki. 1989. temporal reference in english and japanese. austin: university of texas dissertation. ogihara, toshiyuki. 1996. tense, attitudes, and scope. dordrecht: kluwer. o’grady, william. 1997. syntactic development. chicago: the university of chicago press. o’grady, william, miseon lee & miho choo. 2003. a subject-object asymmetry in the acquisition of relative clauses in korean as a second languages. studies in second languages acquisition 25(3). 433–448. https://doi.org/10.1017/s0272263103000172. von stechow, arnim. 2009. tenses in compositional semantics. in wolfgang klein & ping li (eds.), the expression of time, 129–166. berlin: de gruyter. https://doi.org/10. 1515/9783110199031.129. stowell, tim. 1996. the phrase structure of tense. in johan rooryck & laurie zaring (eds.), phrase structure and the lexicon, 277–291. berlin: springer. trueswell, john c. & michael k. tanenhaus. 1991. tense, temporal context and syntactic ambiguity resolution. language and cognitive processes 6(4). 303–338. https://doi.org/ 10.1080/01690969108406946. proceedings of elm 2: 1-12, 2023 giuliano armenante, vera hohaus and britta stolterfoht: transparency in the processing of temporal ambiguity: the case of embedded tense. 12 https://doi.org/10.3765/elm https://www.elm-conference.net/ comparing global and local accommodation: rating and response time data alexander göbel & florian schwarz* abstract. this paper addresses the question to what extent global and local accommodation should be viewed as sharing the same underlying mechanism or whether they are distinct processes that only happen to share the same label. we present offline rating data and response times from a mouse-tracking experiment that directly compared global and local accommodation for five different triggers. the results show that globally accommodating a presupposition led to a larger decrease in acceptance than locally accommodating, and that response times for local accommodation were overall faster. while we take the results to be inconclusive with regard to the question about the underlying mechanism, we conjecture that the contexts tested here were more favorable for local accommodation, and that hence investigating how different contexts affect the relative ease of accommodation type is a promising avenue for future research. keywords. presupposition; accommodation; mouse-tracking 1. introduction. presuppositions are traditionally viewed as content that is taken for granted or backgrounded. however, this status does not map with full generality to a need for presuppositions to be directly satisfied in the preceding context. for instance, the utterance in (1) is intuitively acceptable even without the addressee having been aware of its presupposition, i.e. presuppositions can be accommodated. (1) gordon stopped smoking. ⇝ gordon was smoking before the presupposition literature distinguishes different types of accommodation, most commonly between global and local accommodation. global accommodation describes cases like (1) where the presupposition is interpreted at the root level and becomes part of the speaker’s commitment. local accommodation, on the other hand, can only occur under embedding, when the presupposition is interpreted under the scope of an operator and practically cancelled at the global level, as illustrated in (2). (2) caitlin’s birthday is next week, but i don’t know whether isabelle is planning a surprise party for her. if caitlin realizes beforehand that isabelle is planning a surprise party, then isabelle will probably be very disappointed. ⇝ if isabelle is planning a surprise party and caitlin realizes it, then isabelle will probably be very disappointed the question we want to address in this paper is to what extent the shared label of ‘accommodation’ for (1) and (2) should be taken as indicative of a shared underlying mechanism. we present *we want to thank the experimental study of meaning lab and audiences at hsp 2022 and elm2 for discussion. authors: alexander göbel, mcgill university (alexander.gobel@mcgill.ca) & florian schwarz, university of pennsylvania (florians@sas.upenn.edu). proceedings of elm 2: 95-103, 2023 c©2023 alexander göbel and florian schwarz published by the lsa with permission of the author(s) under a cc by license. 95 https://doi.org/10.3765/elm https://www.elm-conference.net/ data from speeded acceptability judgments from a mouse-tracking experiment directly comparing global and local accommodation to see if they pattern similarly or not.1 the paper is structured as follows. section 2 provides additional background on global and local accommodation. section 3 presents the experiment and section 4 concludes with the general discussion. 2. background. 2.1. theoretical. although global and local accommodation are central concepts in semantic theory, the question of how they relate to each other is rarely discussed. one account of a unified treatment of accommodation comes from krahmer & beaver (2001). the authors propose to model different types of accommodation via an a operator that turns presupposed content into asserted content and can be inserted at different positions at lf. applied to cases of global and local accommodation, this operator would then appear either at the highest level or within the scope of the relevant embedding operator, reducing the difference to one of syntactic position. in contrast, von fintel (2008) argues on conceptual grounds that global and local accommodation should be treated as separate mechanisms. his argument is that global accommodation is a hearer’s cooperative reaction to a deficient context that does not match the requirements of an utterance, in response to which the hearer adjusts the context to prevent the conversation from crashing. local accommodation, on the other hand, concerns changing the meaning of an utterance to be compatible with a context that is at odds with a global interpretation of the respective presupposition. on this view, global accommodation may then be considered more of a pragmatic process, whereas local accommodation is more semantic. lastly, klinedinst (2016) proposes a mixed account that is relativized to the type of presupposition trigger under consideration, namely whether a trigger entails its presupposition in addition to presupposing it (e.g. discover) or not (e.g. regret). by virtue of entailing its presupposition, a trigger of this type will allow to be easily accommodated and simply be treated like asserted content regardless of the type of accommodation. in contrast, triggers that do not entail their presupposition are argued to have different sources of difficulty for global and local accommodation. globally accommodating such a trigger requires adjusting a deficient context, as in von fintel’s (2008) view, whereas local accommodation may cause difficulty because the trigger is semantically idle its presupposition gets cancelled such that it no longer contributes anything, resulting in a type of redundancy. while trigger variation is not the main focus of the present study, we included different triggers to explore klinedinst’s (2016) account, see below. 2.2. experimental. regarding previous experimental investigations of presupposition accommodation, the majority of studies has focused on global accommodation, with comparatively few examining local accommodation and to our knowledge none directly comparing the two. the consensus for global accommodation appears to be that there is a robust cost across various triggers and paradigms (see schwarz 2019). a similar picture emerges for local accommodation, although based on a smaller range of evidence. chemla & bott (2013) showed that for negation in a truthvalue judgment study, deriving a locally accommodated interpretation took longer compared to 1we focus on the acceptability judgments here since the mouse-tracking data was inconclusive, but a summary plot can be found in the appendix. proceedings of elm 2: 95-103, 2023 alexander göbel and florian schwarz: comparing global and local accommodation. 96 https://doi.org/10.3765/elm https://www.elm-conference.net/ an interpretation where the presupposition projects. in line with this finding, romoli & schwarz (2015) report slower response times for local accommodation in the covered box paradigm. local accommodation, like global accommodation, thus seems to be a costly process. the question we aim to address in the following experiment is whether one type may be more costly than the other, as such potential differences in cost provide potential evidence for viewing them as separate mechanisms. 3. experiment. 3.1. materials & design. the experiment used stimuli modeled after the suspension contexts from abusch (2010) as in (2) above and similar to materials in mandelkern et al. (2019), but in short dialogues to approximate a conversational environment, shown in (3): (3) sample item a. unembedded (=global accommodation) + satisfaction (=psp met) a: linda loves traveling, and last year she went to vietnam. b: she went to vietnam again this year , so she probably picked up some vietnamese already. b. embedded (=local accommodation) + satisfaction (=psp met) a: linda loves traveling. b: yeah last year she went to vietnam... if she went to vietnam again this year , then she probably picked up some vietnamese already. c. unembedded (=global accommodation) + ignorance (=psp unmet) a: linda loves traveling, but i don’t know whether she’s been to vietnam before. b: she went to vietnam again this year , so she probably picked up some vietnamese already. d. embedded (=local accommodation) + ignorance (=psp unmet) a: linda loves traveling. b: yeah though i don’t know whether she’s been to vietnam before... if she went to vietnam again this year , then she probably picked up some vietnamese already. each dialogue consisted of four clauses. the third clause was the target clause, highlighted here for the reader’s convenience by framing (not shown to participants), and contained the presupposition trigger. as a first factor, we manipulated accommodation type by either using a root clause for the target connected to the following clause by so (global) or by having the target clause be the antecedent of a conditional followed by its consequent (local). additionally, the second clause in the overall dialogue was either part of the first interlocutor’s speech (global) or part of the second interlocutor’s (local) in order to sidestep a potential global accommodation interpretation of the target clause in the local psp unmet condition. as a second factor, the second clause of each dialogue either satisfied the relevant presupposition (psp met) or not (psp unmet). there were 32 item sets, distributed in a latin-square design, with four (types of) triggers proceedings of elm 2: 95-103, 2023 alexander göbel and florian schwarz: comparing global and local accommodation. 97 https://doi.org/10.3765/elm https://www.elm-conference.net/ evenly split: again, even, still, and factive verbs. factives were additionally split into a cognitive factive (discover) and an emotive factive (regret). 16 non-presuppositional filler items sharing features with the experimental items were added as controls, with half constructed to be rated good and half as bad. the full list of items can be found in the osf repository associated with this publication (https://osf.io/x4yad/) as well as through the experiment link in the subsection below. 3.2. procedure. the experiment was implemented through pc-ibex (zehr & schwarz 2018). each trial began with a button displayed at the center bottom of the screen and large thumbs-up and -down icons at the top left and right respectively. button click started a character-by-character unfolding of the text (at 60ms/char). 500ms before the end of the target clause, participants were prompted via the appearance of a ’green light’ image to quickly indicate acceptability of the discourse so far by moving their cursor to one of the icons as the rest of the line continued to unfold. the initial choice had to happen within 2 seconds. error messages were displayed if the cursor was moved too early or did not reach an icon within the time limit. upon selection, the final clause unfolded, and participants could adjust their acceptability decision if so desired. a demo link can be found here: https://farm.pcibex.net/r/caxwxc/. 3.3. participants. 75 students from the university of pennsylvania were recruited and received course credit as compensation. 7 participants were excluded due to the difference in acceptance rate of good and bad catch fillers being less than 1 3 , leaving 68 participants for data analysis. 3.4. predictions. on the view that global and local accommodation originate from the same underlying mechanism, there is no immediate reason for them not to pattern together, barring other factors or assumptions. that is, global and local accommodation should show a comparable cost, whatever that cost may be. in contrast, evidence for a difference in costs would prima facie favor accounts that assume distinct mechanisms at play, as unified accounts don’t come with an inherent explanation of such differences. (conversely, not finding any differences in cost does not necessarily speak against distinct mechanism accounts, as the measure at hand could simply fail to differentiate costs, or qualitatively different costs need not map onto quantitative differences in the task.) finally, if only triggers that entail their presupposition have a unified mechanism whereas triggers that do not entail their presupposition behave differently for global and local accommodation, we should see no difference in accommodation cost for discover as an entailing trigger and a potential difference in cost for regret, even, and again as non-entailing triggers (see sudo 2012, djärv et al. 2017). for still, there are no prior claims about its entailment status such that it will be put aside in this regard. 3.5. results. responses. the average acceptance rate by condition for responses that did not change after the target clause was presented is shown in figure 1. the first thing to note is that unmet conditions have a much lower acceptance rate than met conditions, indicating a general accommodation cost. this result is shown in our analysis (mixed effects logistic regression with sum coding) as a significant effect of context (z = 15.22, p < .001***). additionally, this accommodation cost is smaller for local accommodation than for global accommodation, as shown by a significant interaction between context and accommodation type in our model (z = 3.05, p < .01**). proceedings of elm 2: 95-103, 2023 alexander göbel and florian schwarz: comparing global and local accommodation. 98 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: response choices by condition with standard errors. looking at acceptance rates by individual triggers shown in figure 2, we see that local accommodation is numerically clearly easier than global accommodation for all triggers but discover, which also has the smallest accommodation cost overall. figure 2: response choices by condition by trigger with standard errors. response times. the average response times per condition, split by (initial) response choice, are shown in figure 3. trials where the cursor was moved before the signal was given or did not proceedings of elm 2: 95-103, 2023 alexander göbel and florian schwarz: comparing global and local accommodation. 99 https://doi.org/10.3765/elm https://www.elm-conference.net/ reach one of the icons to indicate a response within the time limit were excluded (265 or 12% of trials). looking at acceptances first, we see overall faster response times for local accommodation than global accommodation, both when the presupposition was met in the context and when it was unmet. for rejections, the data for met conditions should be taken with a grain of salt due to the small sample size (since less than 20% of responses were rejections here), but for unmet conditions local accommodation is again numerically faster than global accommodation. interestingly, when we only look at unmet conditions, we can see that acceptances for local accommodation were faster than rejections, whereas response choice did not substantially affect global accommodation. this effect is supported by a significant interaction between accommodation type and response choice (z = 2.08, p < .05*) in a mixed effects linear regression model restricted to data for unmet conditions. figure 3: response times by condition with standard errors. 3.6. discussion. the experiment provides evidence that when directly compared local accommodation is easier than global accommodation in our materials. this pattern was reflected in the acceptance rates as well as response times: first, when the respective presupposition was not met, acceptance responses for global accommodation showed a larger decrease than local accommodation; second, participants were faster to accept than to reject trials in the unmet condition for local accommodation, whereas latencies for global accommodation remained about the same. taken at face value, the data would thus be more in line with a view on which global and local accommodation have different underlying mechanisms despite the shared label. however, it is worth taking into account the materials used to investigate our research question. the comparison that allowed us to measure accommodation difficulty was between a prior sentence directly satisfying the presupposition in question and an explicit statement of ignorance regarding the presuppositional content. this comparison contrasts with prior investigations on both global accommodation, where the relevant difference is often between explicit satisfaction and a neutral proceedings of elm 2: 95-103, 2023 alexander göbel and florian schwarz: comparing global and local accommodation. 100 https://doi.org/10.3765/elm https://www.elm-conference.net/ statement that is agnostic (e.g. tiemann et al. 2015), and local accommodation, where sentences are often presented out of context. these changes might have shifted the scales in favor of local accommodation, leaving open what the more general conclusion should be here. we will pick up this issue in the next section. additionally, it is also worth looking at the distribution of the individual triggers we tested. while the aggregated pattern of the results was also present for again, still, even, and regret, discover diverged from the other triggers by showing comparable costs for both accommodation types. notably, discover was the only trigger in the set that had been categorized as entailing its presupposition. such a split, if general, would thus support klinedinst’s (2016) account, according to which triggers that entail their presupposition allow a unified treatment while triggers that do not entail their presupposition may differ in the mechanisms for global and local accommodation. in addition to how the results bear on our main question about the relationship between global and local accommodation, another interesting aspect is the relative ease of local accommodation. as laid out in section 2.2, this pattern is unexpected based on prior studies. possible explanations for this difference will be discussed in the next section. 4. general discussion. this paper presented evidence from speeded acceptability ratings for global and local accommodation differing in their underlying mechanism, but only for triggers that do not entail their presupposition. interestingly, for those triggers, local accommodation was easier than global accommodation. this finding raises the question of why prior studies such as chemla & bott (2013) and romoli & schwarz (2015) found local accommodation to come with a clear cost, in line with its characterization as a last resort strategy in the formal theoretical literature. one obvious difference lies in the methodologies. chemla & bott (2013) used a truth-value judgment task and romoli & schwarz (2015) a covered box paradigm. while their argument is based on response time data from these methodologies, their tasks did not put participants under any time pressure to give their response. in contrast, in our experiment participants had to initiate their cursor movement while the target sentence was still unfolding and had little time once it was complete to indicate their response. the current data may thus be a more realistic representation of the on-line processing of local accommodation: longer response times in untimed studies might have been a reflection of participants who ultimately wind up with a local accommodation interpretation, but only arrive at it reluctantly. in contrast, by putting participants under time pressure in our task, such participants might have been more likely to simply reject the sentence in the case of local accommodation if this interpretation was not available to them right away. redoing the studies from chemla & bott and romoli & schwarz under similar conditions might therefore be an interesting next step forward.2 however, an alternative explanation for why local accommodation was comparatively easy here, which we think is more likely to play a larger role here (though it’s not necessarily incompatible with the previous possibility), relates to the role of context. both chemla & bott (2013) and 2a second notable difference between these two studies and the present one is the type of embedding. both chemla & bott and romoli & schwarz used negation as the embedding operator, whereas here we used the antecedent of a conditional. however, as far as we are aware, there is no discussion of the difference between embedding operators affecting the rate of local accommodation, so there is no prima facie reason why negation should behave differently. proceedings of elm 2: 95-103, 2023 alexander göbel and florian schwarz: comparing global and local accommodation. 101 https://doi.org/10.3765/elm https://www.elm-conference.net/ romoli & schwarz (2015) investigated sentences in isolation without prior context. in contrast, our stimuli were both more naturalistic by using dialogues to approximate actual conversations and made the status of the relevant presupposition explicit. the characterization of local accommodation as a last resort strategy that incurs processing cost may thus be modulated by whether or not it is contextually motivated. as shown in the response choices, local accommodation still comes with a cost and leads to lower acceptance rates. however, making the choice to accept the sentence, indicating local accommodation, may not require much processing cost in itself, if the context supports such an interpretation. conversely, the relative ease of local accommodation compared to global accommodation may have also been due to features of the context. the standard characterization of global accommodation as a cooperative rescue strategy concerns cases when there is no information about whether a presupposition is true or not. in the cases tested here, the first speaker expressed explicit ignorance regarding the presupposition of the second speaker’s utterance, which might have biased against taking it to be true. overcoming this bias might have then stacked the cards against global accommodation. in contrast, since local accommodation does not involve the speaker committing to the truth of the presupposition in the first place, and they themselves had expressed their own ignorance in the previous clause, this additional hurdle did not exist in this case. from this perspective, caution is warranted in interpreting the present results as providing a fully general comparison between accommodation types, insofar as the contexts may have been more beneficial for local accommodation than global accommodation. rather, it seems necessary to gather data across different contexts before being able to conclusively answer the question of whether one type of accommodation is easier than the other (and if so, which). one concrete modification could be to use contexts that leave the truth of the presupposition open without expressing explicit ignorance, as is commonly used in the study of global accommodation. we might expect global accommodation to become easier in this case in the absence of bias, and local accommodation to become harder since the context no longer enforces it. paying closer attention to the context might in turn lead to more insights into the nature of the different accommodation types and their underlying mechanism. appendix. the mouse-tracking data was analyzed using the r package ‘mousetrap’. figure 4 shows the normalized average trajectory by condition. the statistical analysis through the package did not yield any clear results. since the experiment was run online, we relied on written instructions for participants to move their cursor in a less direct manner toward the response icon. additionally, we showed an upward arrow between two bars to encourage participants moving upward before going toward the response icons. references abusch, dorit. 2010. presupposition triggering from alternatives. journal of semantics 27. 37–80. chemla, emmanuel & lewis bott. 2013. processing presuppositions: dynamic semantics vs pragmatic enrichment. language and cognitive processes 28. 241–260. djärv, kajsa, jeremy zehr & florian schwarz. 2017. cognitive vs. emotive factives: an experimental differentiation. in robert truswell, chris cummins, caroline heycock, brian rabern & hannah rohde (eds.), proceedings of sinn und bedeutung, vol. 21, 161–176. proceedings of elm 2: 95-103, 2023 alexander göbel and florian schwarz: comparing global and local accommodation. 102 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 4: mouse trajectories by condition and response choice. von fintel, kai. 2008. what is presupposition accommodation, again? philosophical perspectives 22. 137–170. klinedinst, nathan. 2016. two types of semantic presuppositions. in keith allan, alessandro capone & istvan kecskes (eds.), pragmemes and theories of language use, perspectives in pragmatics, philosophy & psychology 9, 601–624. springer. krahmer, emiel & david beaver. 2001. a partial account of presupposition projection. journal of logic, language and information 10. 147–182. mandelkern, matthew, jeremy zehr, jacopo romoli & florian schwarz. 2019. we’ve discovered that projection across conjunction is asymmetric (and it is!). linguistics and philosophy online. 1–42. https://doi-org.silk.library.umass.edu/10.1007/ s10988-019-09276-5. romoli, jacopo & florian schwarz. 2015. an experimental comparison between presuppositions and indirect scalar implicatures. in florian schwarz (ed.), experimental work on presuppositions, 215–240. springer. schwarz, florian. 2019. presuppositions, projection, and accommodation theoretical issues and experimental approaches. in chris cummins & napoleon katsos (eds.), handbook of experimental semantics and pragmatics, 83–113. oxford: oxford university press. sudo, yasutada. 2012. on the semantics of phi features on pronouns: mit dissertation. tiemann, sonja, mareike kirsten, sigrid beck, ingo hertrich & bettina rolke. 2015. presupposition processing and accommodation: an experiment on wieder (‘again’) and consequences for other triggers. in florian schwarz (ed.), experimental perspectives on presuppositions, 39–65. springer. zehr, jeremy & florian schwarz. 2018. penncontroller for internet based experiments (ibex). https://doi.org/10.17605/osf.io/md832. proceedings of elm 2: 95-103, 2023 alexander göbel and florian schwarz: comparing global and local accommodation. 103 https://doi.org/10.3765/elm https://www.elm-conference.net/ effects of instruction on semantic and pragmatic judgment tasks ziling zhu & dorothy ahn* abstract. sentence judgment tasks are used often in linguistics studies. however, there is no consensus on how significant the effect of instruction is in such tasks: some argue that instruction is trivial, while others argue that it affects the way participants respond. in this study, we investigate different keywords used in sentence judgment tasks and determine which keyword best teases apart speakers’ response to semantically and pragmatically licit and illicit sentences. we test this in english and mandarin, exploring the possibility of cross-linguistic variation on how speakers respond to different keywords. our results show that the common keywords used in semantic and pragmatic judgment tasks such as ‘natural’ do distinguish semantic and pragmatic violations for english speakers, but that the common mandarin translations of these words fail to distinguish between the two types of violations. our results highlight the need for languageand study-specific norming procedures in sentence judgment tasks. keywords. sentence judgment tasks; instruction type; semantics; pragmatics; psycholinguistics; cross-linguistic variation 1. introduction. experimental linguistic work is defined by its design, procedures, and statistical analysis (kirk 2012, myers 2017). there have recently been more discussions on how to optimize procedures for sentence judgment tasks, featuring two considerations: the instruction (schütze 2005; a.o.) and the response scale (schütze & sprouse 2013; a.o.). instruction conveys what the researchers ask the participants to do, and response scale determines how the participants communicates back to the researchers. in this project, we focus on the role of instruction in a commonly used experimental paradigm in linguistics: sentence judgment tasks. empirical evidence diverges as to whether instruction variation matters in sentence judgment tasks. on the one hand, instruction has been claimed to be trivial in morphological (aronoff & schvaneveldt 1978) and syntactic (cowart 1997) judgment tasks. on the other hand, the effect of instruction was reported in some other syntactic studies, where different keywords lead to significantly different judgments from the participants (maclay & sleator 1960, beltrama & xiang 2016). methodological studies on pragmatic judgment tasks have been conducted in veenstra & katsos (2018), but the role of instruction variation has received scarce attention. this study fills the research gap for experimental semantics and pragmatics, revealing that instruction is a significant factor in identifying and distinguishing between semantic and pragmatic violations in sentence judgment tasks. furthermore, we show that english and mandarin speakers respond differently to keywords in the instructions, highlighting the need for language and studyspecific norming procedures. this paper is organized as follows: section 2 introduces the background, discussing the contrasting claims on whether instruction type is a significant factor or not in sentence judgment tasks. *we would like to thank audience at elm2 and rutgers 2022 open house for their helpful discussions. funding was generously provided by meaning across languages (mal) lab of rutgers university. special thanks go to chaoyi chen and joseph casillas for techinical support on statistics. authors: ziling zhu, rutgers university (ziling.zhu@rutgers.edu) & dorothy ahn, rutgers university (dorothy.ahn@rutgers.edu). proceedings of elm 2: 322-330, 2023 c©2023 ziling zhu and dorothy ahn published by the lsa with permission of the author(s) under a cc by license. 322 https://doi.org/10.3765/elm https://www.elm-conference.net/ in section 3, we report on the results of the study we ran, which suggest that a) instruction type is significant for identifying semantic and pragmatic violations and b) english and mandarin speakers respond differently to keywords that are often used as translations of each other in empirical studies. section 4 concludes with a discussion of implications and remaining questions. 2. background. instruction variation, as a procedural factor of linguistic experiments, was first investigated by hill (1961). in hill’s study, ten subjects were instructed to “reject any sentences which were ungrammatical, and to accept those which were grammatical”, with no definition of (un)grammaticality. the results showed that participants lacked a consistent understanding of this concept. however, since cowart’s (1997) experiment, to be reviewed next, it has been assumed in experimental linguistics that “the exact nature of instructions matters relatively little” (schütze & sprouse 2013), although systematic empirical investigations are not yet in place. in answering this methodological question, as to whether instruction variation is significant in linguistic experiments, two claims have been put forward in the literature. 2.1. instruction variation is trivial. instruction variation was claimed to be trivial for morphological (aronoff & schvaneveldt 1978) and syntactic (cowart 1997, schütze & sprouse 2013) judgment tasks. we review the following studies that support this view. to test the morphological productivity of word formation rule with affixes, aronoff & schvaneveldt (1978) designed some possible but non-occurring english words and asked participants to evaluate them. crucially, they manipulated the instruction as in (1). (1) a. is the item in your vocabulary? b. is the item an english word? c. is the item a meaningful word? each group of participants only saw one question. the percentage of affirmative responses were calculated and reported. the researchers reported that although the percentages vary as the instructions vary (42% for (1-a), 50% for (1-b), 54% for (1-c)), the instruction variation has no significant effect on how morphology influences participants’ responses. in the field of syntax, cowart (1997) claims that the nature of instructions matters little. two types of instructions, termed as ‘intuitive’ and ‘prescriptive’ by cowart, were used in this study. the intuitive instructions in (2-a) highlight the participants’ perception of these sentences, namely whether the stimuli are fully normal/very odd, or somewhere between the extremes. in contrast, the prescriptive instructions in (2-b) invoke a prescriptive view, pointing the participants to school grammar and language authorities. (2) a. intuitive instructions: please read each of the sentences listed below. for each sentence, we would like you to indicate your reaction to the sentence. mark your response sheet a, b, c, or d. use (a) for sentences that seem fully normal, and understandable to you. use (d) for sentences that seem very odd, awkward, or difficult for you to understand. (note: do not use “e”.) if your feelings about the sentence are somewhere between these extremes, use one of the middle responses, b or c. there are no “right” or “wrong” answers. please base your responses solely on your gut reaction, not on rules you may have learned about what is “proper” or “correct” english. b. prescriptive instructions: please read each of the sentences listed below. for each proceedings of elm 2: 322-330, 2023 ziling zhu and dorothy ahn: effects of instruction on semantic and pragmatic judgment tasks. 323 https://doi.org/10.3765/elm https://www.elm-conference.net/ sentence, we would like you to indicate whether or not you think the sentence is a wellformed, grammatical sentence of english. suppose this sentence were included in a term paper submitted for a 400-level english course that is taken only by english majors; would you expect the professor to accept this sentence? mark your response sheet a, b, c, or d. use (a) for sentences that seem completely grammatical and well-formed. use (d) for sentences that you are sure would not be regarded as grammatical english by any appropriately trained person. (note: do not use “e”.) if your judgment about the sentence is somewhere between these extremes, use one of the middle responses, b or c. use b for sentences you think probably would be accepted but you are not completely sure. use c for sentences you think probably would not be accepted. participants were asked to evaluate sentences with local/remote antecedents in different syntactic structures. cowart reports no linguistically meaningful influences of the two types of instructions. 2.2. instruction variation is significant. in contrast to the previous view, task effect has been observed in other linguistic judgment tasks (maclay & sleator 1960, beltrama & xiang 2016), suggesting that instruction variation could be a significant factor in sentence judgment tasks. to explore the nature of sentence judgment tasks in linguistics, maclay & sleator (1960) devised six groups of sentence stimuli differing in whether they are ‘grammatical’, ‘meaningful’, and ‘ordinary’. the instructions vary accordingly, as in (3). (3) a. do these words form a grammatical english sentence? b. do these words form a meaningful english sentence? c. do these words form an ordinary english sentence? participants were asked to answer yes or no, and the proportion of affirmative responses were calculated. the researchers observed that judgments of syntactic well-formedness and those of semantic meaningfulness are independent from each other. for stimuli with semantic violations (grammatical but not meaningful), 42% of the participants judged them to be grammatical, while only 7% judged them to be meaningful. in contrast, for stimuli with syntactic violations (ungrammatical but meaningful), 34% of the participants judged them to be grammatical, while 52% judged them to be meaningful. participants systematically rejected syntactic violations more with the ‘grammatical’ instruction in (3-a), and semantic violations more with the ‘meaningful’ instruction in (3-b). in other words, the instructions ‘grammatical’ and ‘meaningful’ could systematically tease apart participants’ response to syntactic and semantic violations, respectively. similar observations have been made in a syntactic study by beltrama & xiang (2016).1 to examine whether intrusive resumptive pronouns can rescue island violations, they designed an acceptability task and a comprehensibility task, with instructions in (4-a) and (4-b) respectively. (4) a. how acceptable is the [target sentence]? please make your judgments based on how good the [target] sentence sounds in english given the context it is in. b. we want you to judge these sentences based on how easy they are for you to understand. whereas resumptive pronouns do not improve island violations in the acceptability task, such rescuing effect was found in the comprehensibility task. this task effect crucially arises from the 1we thank troy messick for pointing us to this study. proceedings of elm 2: 322-330, 2023 ziling zhu and dorothy ahn: effects of instruction on semantic and pragmatic judgment tasks. 324 https://doi.org/10.3765/elm https://www.elm-conference.net/ instruction variation in (4), suggesting that instruction is in fact significant in guiding participants’ responses in syntactic sentence judgment tasks. we fill two research gaps in this project. first, we explore whether instruction is a significant factor in semantic and pragmatic sentence judgment tasks. as reviewed above, most linguistic studies that examine instruction variation in sentence judgment tasks have focused on morphology and syntax. second, we extend our research question to make a cross-linguistic comparison. specifically, we ask whether english and mandarin speakers respond differently to keywords in the instructions that are often assumed to be comparable to each other in empirical studies (hara et al. 2014, xue et al. 2020, law & syrett 2017). 3. experiment. to investigate the effects of instruction in semantic and pragmatic sentence judgment tasks, we compared participants’ responses to different instructions against the same set of sentence stimuli. 3.1. stimuli. we chose four commonly used instructions in english sentence judgment tasks, shown in (5), varying in the key adjectives. (5) a. does this sound natural to you? b. does this sound acceptable to you? c. does this sound grammatical to you? d. how likely is it for a native speaker to say this? in order to test for language-specific effects, we also created a mandarin version of the english instructions as in (6), using words commonly used as translations of those found in (5). for example, ‘ziran (natural)’ was used in hara et al. (2014), ‘ke jieshou (acceptable)’ in xue et al. (2020) and law & syrett (2017), among many others. (6) a. yixia following neirong contents ting-qilai hear-impression ziran natural ma? q-part? ‘do the following contents sound natural?’ b. yixia following neirong contents ting-qilai hear-impression fuhe fit yufa grammar ma? q-part? ‘do the following contents sound grammatical?’ c. yixia following neirong contents ting-qilai hear-impression ke can jieshou accept ma? q-part? ‘do the following contents sound acceptable?’ d. nin you renwei think muyu native.language wei be hanyu mandarin de gen ren, person, you have duo-da how-big keneng possibility shuo-chu say-out yixia following neirong? contents? ‘how likely do you think is it for a native speaker of mandarin to say the following contents?’ a total of 24 syntactically well-formed sentences were tested as the stimuli. we grouped them into three categories based on their semantic and pragmatic felicitousness. the first group contains 8 semantically odd stimuli involving lexical contradictions (7-a)–(7-c), logical contradictions proceedings of elm 2: 322-330, 2023 ziling zhu and dorothy ahn: effects of instruction on semantic and pragmatic judgment tasks. 325 https://doi.org/10.3765/elm https://www.elm-conference.net/ (7-d),(7-e), and thematic mismatch (7-f)–(7-h). our categorization of these as semantic violations is based on a few assumptions of what falls under semantic knowledge. first, we assume that world knowledge based on lexical meaning (e.g. a bachelor is unmarried) constitutes semantic knowledge of a word. this kind of lexical contradiction is what is often called ‘semantic violations’ in eeg studies, for example. second, we assume that logical relations between propositions are part of the semantic knowledge that a speaker has. note that world-knowledge and logical relations are quite different from each other. for now, we group these together under ‘semantic violations’, but it would be good to test whether the two kinds of violations elicit different responses. (7) a. jake is a married bachelor. b. jasmine talked silently. c. bantee’s yellow hat is blue. d. it is raining and not raining outside. e. i’m lying when i say this sentence is true. f. zhangsan smelled 3 o’clock. g. i bought a lamp, and the lamp is drinking water. h. i washed brightness. the second group consisted of 8 pragmatically odd stimuli with redundant information, including direct repetition of information (8-a)–(8-d) and repetition of scalar implicature (8-e)–(8-h). our categorization of pragmatic violations is based on two factors. first, the sentences do not meet our definition of semantic violations: they are neither logically nor lexically contradictory. second, the sentences violate certain pragmatically-motivated constraints such as the gricean maxims. redundant information, for example, is a violation of the quantity maxim, while repetition of scalar implicature is a violation of the manner maxim. (8) a. yuki arrived. yuki sat down. yuki turned on her laptop. b. mimi jumped onto the bed. mimi cried. mimi decided to sleep. c. it rained yesterday when it rained yesterday. d. carolyn went to the park when she went to the park. e. not only are all the students sad, some of them are sad. f. all the students passed the exam and some of them passed. g. three engineers came to work today and two engineers came to work today. h. becky and vera went to the party and becky went to the party. the third group contained 8 neutral stimuli with no identifiable semantic or pragmatic violations as defined above. the sentences presented under this categorization are shown in (9). (9) a. yuki decided to go to school today. b. mason thinks it’s raining outside. c. anya was drinking water, because she was thirsty. d. i bought a beautiful hat three days ago. the hat was yellow. e. the students turned off their laptops and went outisde. f. yesterday it was sunny in toronto. g. a bachelor walked up to us and introduced himself. proceedings of elm 2: 322-330, 2023 ziling zhu and dorothy ahn: effects of instruction on semantic and pragmatic judgment tasks. 326 https://doi.org/10.3765/elm https://www.elm-conference.net/ h. i have a daughter and a son. they are nice to each other. 3.2. participants and procedure. we recruited 81 native english speakers and 81 native mandarin speakers (18–64; gender-balanced) via prolific. participants were redirected to a qualtrics survey, where they were asked to first provide some demographic and language background information and then complete the sentence judgment task. participants were compensated $2-3 for their time. the study was designed to be between-subject, so that each participant would only see one instruction type for all 24 test items. participants were presented with the sentence stimuli (randomized in order) one at a time and were asked to respond on a 7-point likert scale based on the instruction they saw, as in fig. 1. figure 1: sample question (‘natural’ condition with a lexical contradiction) we collected the ratings for each sentence as the dependent variable and tested for each language whether stimuli group (neutral, semantically odd, pragmatically odd) and/or instruction type (natural, acceptable, grammatical, likely) lead to significant rating differences. 3.3. predictions. if instruction variation is trivial for semantic and pragmatic judgment tasks, we would predict that instruction type would not change the rating results for each test sentence. instead, participants would rate the sentences based on their respective standards. if instruction variation is significant, however, different instructions would lead to different ratings of the same stimuli. if certain keywords are more likely to prime judgments based on certain violations, we would also expect the contrasts between conditions to be consistent across speakers. finally, if the (in)significance of instruction variation has no cross-linguistic difference, then native english speakers and native mandarin speakers would show similar contrasts across different instruction types. 3.4. results. we fit a cumulative link mixed model in r to compare ratings in different conditions (fig. 2). we first observed a cross-linguistic variation for instruction type: for english, the results showed a main effect of stimuli group (p < 0.001), instruction type (p < 0.001), and a significant interaction (p < 0.001); for mandarin, we only found a main effect of stimuli group (p < 0.001), but not instruction type (p > 0.1), and no significant interaction (p > 0.1). this suggests that, while english speakers use different instructions to tease apart different proceedings of elm 2: 322-330, 2023 ziling zhu and dorothy ahn: effects of instruction on semantic and pragmatic judgment tasks. 327 https://doi.org/10.3765/elm https://www.elm-conference.net/ linguistic violations, mandarin speakers do not make this distinction among the instructions. english mandarin figure 2: ratings as function of stimuli group, grouped by instruction type (n: neutral; p: pragmatically odd; s: semantically odd) across the stimuli groups, all instruction types reliably distinguished between odd and neutral stimuli (p < 0.001) for both english and mandarin. between semantically and pragmatically odd sentences, for english, all instruction types led to significantly different responses except for grammatical (p > 0.1); for mandarin, all instruction types led to significantly different responses (p < 0.001). moreover, the instruction type natural was the most effective in teasing apart the stimuli groups for both english and mandarin. 4. discussion and conclusions. our experiment reveals the significance of instruction type in semantic and pragmatic sentence judgment tasks. first, we confirm the intuitive choice, made by previous researchers, of using ‘natural’ in the instruction (cremers & chemla 2017, zlogar & davidson 2018, hara et al. 2014; a.o.), which seems to distinguish between semantically and pragmatically odd sentences most clearly. second, we highlight the need to include control sentences with standard ratings to evaluate semantic and pragmatic violations more accurately. sprouse et al. (2022) use a set of previously-tested sentences as fillers to calibrate newly collected grammaticality judgments in their syntax study. our preliminary data can serve a similar role in semantic and pragmatic judgment tasks. it is important to note that by ‘teasing apart semantic and pragmatic violations’, we do not mean that there is a clearly delineated diagnostic that determines whether something is a semantic violation or a pragmatic violation. what the stimuli offer instead is a way to compare the rating of a given sentence against the ratings associated with what we know to be logically or pragmatically illicit and indirectly determine the rationale for the participants’ response. for example, the stimuli used in the current study has been used as controls to an independent study looking at participants’ rating of mandarin bridging anaphora (zhu & ahn in prep). comparing participants’ rating of the target stimuli against the controls allowed us to determine whether participants’ rating of bridging anaphora align better with semantically odd sentences or pragmatically odd sentences. proceedings of elm 2: 322-330, 2023 ziling zhu and dorothy ahn: effects of instruction on semantic and pragmatic judgment tasks. 328 https://doi.org/10.3765/elm https://www.elm-conference.net/ the current study also draws attention to cross-linguistic differences in sentence judgment tasks. first, we see that the range of ratings spreads wider in mandarin than in english in general. the central tendency bias, where participants avoid the endpoints of a scale, is considered to be one of the most robust biases found in psychology (stevens 1971). our data from mandarin participants, where neutral sentences are rated at the ceiling, suggest that there might be crosslinguistic variation on this tendency as well. this raises a question on the nature of this tendency, i.e. whether this is a general cognitive bias that applies across languages, or something that develops culture-specifically. another cross-linguistic difference we observe is that instruction type makes a difference in the way participants respond to different stimuli in english, but not in mandarin. hence, language-specific norming studies with control sentences seem crucial in order to effectively compare cross-linguistic judgments. the grouping of the stimuli into pragmatically odd, semantically odd, and neutral sentences is not independently motivated and thus potentially theory-internal. however, our results suggest that the paradigm of sentence judgment tasks can identify at least some distinction between logically illicit sentences (semantically odd) and sentences that are logical but not discourse-natural (pragmatically odd). in the future, a clustering study could help identify and distinguish between subtypes of semantic and pragmatic violations. references aronoff, mark & roger schvaneveldt. 1978. testing morphological productivity. annals of the new york academy of sciences 318(1). 106–114. https://doi.org/10.1111/j.17496632.1978.tb16357.x. beltrama, andrea & ming xiang. 2016. unacceptable but comprehensible: the facilitation effect of resumptive pronouns. glossa: a journal of general linguistics 1(1). https://doi.org/10.5334/gjgl.24. cowart, wayne. 1997. experimental syntax. sage. cremers, alexandre & emmanuel chemla. 2017. experiments on the acceptability and possible readings of questions embedded under emotive-factives. natural language semantics 25(3). 223–261. https://doi.org/10.1007/s11050-017-9135-x. hara, yurie, shigeto kawahara & yuli feng. 2014. the prosody of enhanced bias in mandarin and japanese negative questions. lingua 150. 92–116. https://doi.org/10.1016/j.lingua.2014.07.006. hill, archibald a. 1961. grammaticality. word 17(1). 1–10. https://doi.org/10.1080/00437956.1961.11659742. kirk, roger. 2012. experimental design: procedures for the behavioral sciences. sage publications. law, jess hk & kristen syrett. 2017. experimental evidence for the discourse potential of bare nouns in mandarin. in nels 47: proceedings of the forty-seventh annual meeting of the north east linguistic society, vol. 2, 231–40. maclay, howard & mary d sleator. 1960. responses to language: judgments of grammaticalness. international journal of american linguistics 26(4). 275–282. myers, james. 2017. acceptability judgments. in oxford research encyclopedia of linguistics, proceedings of elm 2: 322-330, 2023 ziling zhu and dorothy ahn: effects of instruction on semantic and pragmatic judgment tasks. 329 https://doi.org/10.3765/elm https://www.elm-conference.net/ https://doi.org/10.1093/acrefore/9780199384655.013.333. schütze, carson t. 2005. thinking about what we are asking speakers to do. in stephan kepser & marga reis (eds.), linguistic evidence: empirical, theoretical, and computational perspectives, 457–485. mouton de gruyter berlin. https://doi.org/10.1515/9783110197549. schütze, carson t & jon sprouse. 2013. judgment data. in robert podesva & devyani sharma (eds.), research methods in linguistics, 27–50. sprouse, jon, troy messick & jonathan david bobaljik. 2022. gender asymmetries in ellipsis: an experimental comparison of markedness and frequency accounts in english. journal of linguistics 58(2). 345–379. https://doi.org/10.1017/s0022226721000323. stevens, stanley s. 1971. issues in psychophysical measurement. psychological review 78(5). 426. veenstra, alma & napoleon katsos. 2018. assessing the comprehension of pragmatic language: sentence judgment tasks. in andreas jucker, klaus schneider & wolfram bublitz (eds.), methods in pragmatics, vol. 10, 1806. walter de gruyter gmbh & co kg. https://doi.org/10.1515/9783110424928. xue, wenting, meichun liu & stephen politzer-ahles. 2020. a study of complement coercion in mandarin chinese: evidence from an acceptability judgment task. in workshop on chinese lexical semantics, 775–784. springer. https://doi.org/10.1007/978-3-030-81197-6 64. zlogar, christina & kathryn davidson. 2018. effects of linguistic context on the acceptability of co-speech gestures. glossa: a journal of general linguistics 3(1). https://doi.org/10.5334/gjgl.438. proceedings of elm 2: 322-330, 2023 ziling zhu and dorothy ahn: effects of instruction on semantic and pragmatic judgment tasks. 330 https://doi.org/10.3765/elm https://www.elm-conference.net/ insensitivity to truth-value in negated sentences: does linear distance matter? sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, & barbara kaup* abstract. affirmative sentences are comprehended more quickly when they are true vs. false but this facilitation is often reduced or absent in negative sentences, yielding a so-called negation-by-truth-value interaction. the reduced sensitivity to truth-value has been attributed to processing difficulties triggered by negation. we investigated whether such difficulties were eased when comprehenders were given more time to process the negator. specifically, we compared negated sentences in which the negator immediately preceded an adjectival predicate vs. occurred earlier in the sentence, separated by several words from the predicate. the results of two sentence-picture matching tasks replicated previous findings of increased processing difficulties in negative vs. affirmative sentences, as well as the negation-by-truth-value interaction. however, we did not find evidence that sensitivity to truth-value was modulated by the distance between the negator and the predicate. our findings suggest that, when sentences are presented in isolation, having more time to process a negator does not confer a measurable comprehension advantage. keywords. negation; linear distance; truth-value; sentence-picture matching; comprehension; german 1. introduction. sentences are usually easier to process when they are true, but this generalization is challenged by negative sentences. this was shown in sentence-picture verification studies, in which participants saw pictures together with affirmative or negative sentences and indicated whether the pictures rendered the sentences true or false. the results showed that affirmative sentences were evaluated more quickly when they were true vs. false, but sensitivity to truth-value was often reduced—or even absent—in negative sentences, giving rise to a “negation-by-truthvalue interaction” (for reviews see kaup & dudschig 2020; carpenter & just 1975). crucially, the negation-by-truth-value interaction was later replicated in tasks without a judgment/verification component, e.g., participants only had to decide whether a pictured object had been mentioned in the sentence (kaup, lüdtke & zwaan 2005; tian, breheny & ferguson 2010). these results suggested that the processing difficulty elicited by negation is a general marker of its comprehension, rather than a by-product of truth-value judgments. this motivated the claim that negative sentences are generally understood in two steps. for example, given the sentence “the package is not wrapped”, comprehenders first represent the counterfactual (or alternate) state-of-affairs expressed by the affirmative proposition (‘the package is wrapped’). later, in a second step, this alternate representation is suppressed, and the actual state-of-affairs is * this research was carried out within project c06 funded by the deutsche forschungsgemeinschaft (dfg, german research foundation) as part of the sfb 1629 “negation in language and beyond”– project number 509468465. authors: sol lago, goethe university frankfurt (sollago@em.uni-frankfurt.de), petra schulz, goethe university frankfurt, esther rinke, goethe university frankfurt, elise oltrogge, goethe university frankfurt, carolin dudschig, university of tübingen, & barbara kaup, university of tübingen. author contributions: sol lago: conceptualization, formal analysis, investigation, software, methodology, writing – original draft, writing – review & editing. petra schulz: conceptualization, methodology, writing – review & editing. esther rinke: conceptualization, methodology, writing – review & editing. elise oltrogge: investigation, software, methodology, writing – review & editing. carolin dudschig: supervision, writing – review & editing. barbara kaup: supervision, writing – review & editing. proceedings of elm 3: 214-223, 2025 c©2025 sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, and barbara kaup published by the lsa with permission of the author(s) under a cc by license. 214 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ represented (kaup, lüdtke & zwaan 2005; kaup, lüdtke & zwaan 2006). because the activation of an alternate interpretation and its suppression are triggered by negative but not by affirmative sentences, 2-step models can explain why negation increases processing time. they can also explain the negation-by-truth-value interaction by proposing that comprehenders create mental simulations of the state-of-affairs described by a sentence. if the sentence is followed by a task to identify a pictured object, responses are faster when the picture matches the simulation created while reading the sentence, and slower when it mismatches this simulation. the lack of processing facilitation for true negative sentences occurs because, at the point of picture identification, the alternate state-of-affairs is still activated, which interferes with identification responses and neutralizes the processing advantage otherwise obtained with true statements. however, later findings suggested that the representation of an alternate state-of-affairs could be diminished or even avoided altogether when negative sentences were pragmatically licensed by context and/or the question-under-discussion was prominent (nieuwland & kuperberg 2008; orenes, beltrán & santamaría 2014; tian, breheny & ferguson 2010; tian, ferguson & breheny 2016; darley, kent & kazanina 2020). for example, tian et al. (2016) used the visual world eyetracking paradigm to demonstrate that when the question-under-discussion was clear to comprehenders, they no longer activated a counterfactual interpretation in english negative sentences. in another visual world study, orenes et al. (2014) showed that english participants could quickly switch their visual attention to the actual state-of-affairs after hearing a sentence like “the figure is not red”, when an alternative interpretation was clearly available, e.g., through a visual context showing only red or green figures. further, an event-related potentials study by nieuwland & kuperberg (2008) demonstrated that brain responses were sensitive to truth-value when negative sentences were preceded by a pragmatically licit linguistic context (e.g., “with proper equipment, scuba-diving isn’t dangerous/*safe”). the findings above indicate that the activation of an alternate interpretation depends on the pragmatic licensing of a negated sentence. the open question is whether non-pragmatic factors may also play a role to help ease the comprehension of negation. one such factor concerns the distance between the negator and its predicate. for example, a negator may appear immediately before an adjectival predicate (as in the example above, “the package is not wrapped”) or farther away, e.g., separated by several words: “it is not true that the package is wrapped”. increased distance might ease the processing of negation either by preventing the activation of an alternate interpretation and/or by facilitating its suppression when the adjectival predicate is encountered. to date, only one study has examined this hypothesis but it found no evidence that processing differences were modulated by the linear position of the negator (dudschig et al. 2019). the study used event-related potentials and measured brain responses to adjectives in true and false german sentences in which the negator occurred either immediately before an adjective or separated by several words, e.g., “ladybirds are not stripy” vs. “it is not true that ladybirds are stripy” —note that the falseness of the sentences was based on world knowledge violations, e.g., about the typical pattern of ladybirds. for negative sentences like “ladybirds are not stripy”, the n400—a negative potential peaking around 400 milliseconds over centro-parietal brain regions—had been previously found to be insensitive to the sentence truth-value (fischler et al. 1983). dudschig et al. (2019) examined whether increasing the distance between the negator and the adjective would yield n400 sensitivity to truth value. the results showed that n400 responses at the adjective were similar in the close and far distance conditions, suggesting that more time to process the negator did not aid comprehension. proceedings of elm 3: 214-223, 2025 sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, and barbara kaup: insensitivity to truth-value in negated sentences: does linear distance matter?. 215 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ however, some methodological aspects make it difficult to directly compare the results of dudschig et al. (2019) with those of previous sentence-picture matching studies (kaup, lüdtke & zwaan 2005; tian, breheny & ferguson 2010). in contrast to dudschig et al. (2019), sentencepicture matching studies measured comprehension after the entire sentence was read and used response times—as opposed to brain responses to one word—as a processing diagnostic of the negation-by-truth-value interaction. to resolve these differences, we adopted the conditions of dudschig et al. (2019) in a sentence-picture matching task in german. like previous studies, we used an implicit version of the task: participants did not have to evaluate the sentences but rather whether a pictured object had been mentioned in the sentence. we conducted two experiments. experiment 1 replicated the negation-by-truth-value interaction reported in previous research by comparing affirmative and negative sentences—the negator in the negative versions immediately preceded the adjectival predicate. experiment 2 compared negative sentences in which the negator was adjacent with the predicate vs. separated by several words, to examine whether more distance—and thus more processing time—would facilitate negation processing, either by preventing the activation of an alternate interpretation and/or by facilitating its suppression later on. if so, we expected to restore participants’ sensitivity to truth-value in far distance negative sentences (but not in close distance sentences), yielding an interaction between truth-value and the distance between the negator and the predicate. 2. methods. 2.1. materials. the critical sentences in experiment 1 consisted of 40 item sets with the structure ‘the noun is {here/not} adjectival predicate’, e.g., “das paket ist hier/nicht eingepackt” (table 1). all items had an affirmative and a negative version, with the negative version featuring the negator “nicht” linearly adjacent to the predicate (i.e., a close distance configuration). the affirmative sentences replaced the negator with the word “hier” (‘here’), such that affirmative and negative sentences had the same number of words. each item set was paired with two pictures depicting either the actual or the alternate state-ofaffairs described in the sentence (e.g., an image of a wrapped vs. an unwrapped package). the pictures were black-and-white drawings, either ai-generated (https://illustroke.com/) or collected from different sources on the web and manually edited if necessary. both the picture and the adjectival predicate (e.g., “eingepackt” vs. “ausgepackt”, ‘wrapped’ vs. ‘unwrapped’) were used to manipulate the state-of-affairs. these two factors were fully crossed to ensure that betweencondition differences were not attributable to differences in the lexical properties of the predicates or in the visual complexity of the images. this resulted in eight latin-square lists, which were collapsed to four in the analysis—since the individual effects of picture and predicate identity were not of theoretical interest for the current study. thus, experiment 1 had a polarity (affirmative/negative) × state-of-affairs (actual/alternate) design. experiment 2 also featured 40 item sets. the (close distance) negative conditions in experiment 1 were retained, but affirmative sentences were replaced by negative sentences in which the distance between the negator and the adjectival predicate was increased by moving the negator to a preceding clause (3 words away from the predicate), e.g., “es stimmt nicht, dass das paket eingepackt ist” (‘it is not true that the package is wrapped’). the identity of the picture and of the predicate were fully crossed, resulting in 8 latin-square lists—collapsed to four in the analysis. thus, experiment 2 featured only negative sentences in a distance (close/far) × state-ofaffairs (actual/alternate) design. proceedings of elm 3: 214-223, 2025 sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, and barbara kaup: insensitivity to truth-value in negated sentences: does linear distance matter?. 216 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ experimental conditions pictures (one picture shown per trial) a. affirmative, actual das paket ist hier ausgepackt. ‘the package is here unwrapped.’ b. affirmative, alternate das paket ist hier eingepackt. ‘the package is here wrapped.’ c. negative close distance, actual das paket ist nicht eingepackt. ‘the package is not wrapped.’ d. negative close distance, alternate das paket ist nicht ausgepackt. ‘the package is not unwrapped.’ e. negative far distance, actual es stimmt nicht, dass das paket eingepackt ist. ‘it is not true that the package is wrapped.’ f. negative far distance, alternate es stimmt nicht, dass das paket ausgepackt ist. ‘it is not true that the packet is unwrapped.’ table 1: sample item set in experiments 1 and 2. conditions (a–d) were used in experiment 1. conditions (c–f) were used in experiment 2. the picture for the actual state-of-affairs is displayed with a dotted line for explanatory purposes only. in the alternative latin-square lists (not shown here), the other image was the target picture, and the adjectival predicate was reversed. 2.2. participants. the participants were self-reported first language speakers of german, who were recruited using the online platform prolific (http://www.prolific.com/). we excluded participants who reported being left-handed, having uncorrected vision or language impairments, or who did not solve at least 80% of the attention checks presented during the experiment (see section 2.3). this resulted in a final sample of 69 participants in experiment 1 (age range: 19–45 years; 29 women, 1 non-binary) and 72 in experiment 2 (age range: 18–44 years; 35 women, 3 non-binary). the experiments were performed in accordance with the declaration of helsinki and the procedure was reviewed and approved by the ethikkommission der deutschen gesellschaft für sprachwissenschaft. all participants provided informed consent to participate in the study. 2.3. procedure. participants completed the sentence-picture matching task online in the testing platform pcibex (zehr & schwarz 2018). sentences were shown word-by-word (soa = 300 ms) and were followed by a picture. participants were instructed to press a key for ‘yes’ when the object shown in the picture appeared in the sentence and ‘no’ when it did not appear. the f and j keys were used for this purpose—their mappings to ‘yes’ and ‘no’ were counterbalanced across participants. the target answer was always ‘yes’ for the experimental sentences. in experiment 1, the picture appeared 400 ms after the sentence offset. in experiment 2, the picture appeared 400 ms after the sentence offset in the close distance conditions, and 100 ms after the sentence offset in the far distance conditions. this ensured that the time elapsed between the presentation of the adjectival predicate and the picture was identical across the close and far distance conditions (i.e., proceedings of elm 3: 214-223, 2025 sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, and barbara kaup: insensitivity to truth-value in negated sentences: does linear distance matter?. 217 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 400 ms). thus, differences between conditions could not be attributed to participants having different amounts of time to plan their answers. the experimental sentences were intermixed with 40 filler sentences. in experiment 1, the filler sentences had the same structure as the experimental sentences but were always affirmative, e.g. “die brille ist jetzt geputzt” (‘the glasses are now cleaned’). to add lexical variation, the word “hier” in the experimental items was replaced with other one-syllable adverbs in the filler items, e.g., “jetzt”/“sehr”/“dort” (‘now’/‘very’/‘there’). all fillers had ‘no’ as a target answer such that ‘yes’ and ‘no’ target responses had a 1:1 ratio across the experiment. in experiment 2, half of the fillers were adapted to start with a preamble comparable to that in the far distance negative sentences, e.g., “es stimmt, dass die brille jetzt geputzt ist” (‘it is true that the glasses now are cleaned’). experimental and filler items were interspersed with 12 attention checks (oppenheimer, meyvis & davidenko 2009). in the attention checks, the sentences were followed by comprehension questions instead of pictures, in order to encourage participants to understand the sentences (e.g., sentence: “the coffee is already cold”; question: “has the coffee cooled down yet?”; response options: yes/no). after 4 practice items, the 92 trials (experimental items, fillers and attention checks) were presented in a randomized manner. an experimental session lasted 10– 15 minutes. 2.4. analysis. raw data were preprocessed manually in order to correct typos and inconsistent demographic responses. the preprocessed data was exported for analysis to r (r development core team 2024). following previous research (kaup, lüdtke & zwaan 2005; tian, breheny & ferguson 2010), the main dependent measure in the analysis was the response time in correctly answered trials. following kaup et al. (2005), we excluded trials with response times shorter than 200 ms or longer than 5000 ms (experiment 1: 0.43–1.3% of trials across conditions; experiment 2: 0.56–2.64% of trials across conditions). following the box-cox procedure (box & cox 1964), response times were reciprocally transformed (–1000/response time). we also analyzed the accuracy of picture responses. accurate responses were coded as 1 and inaccurate responses as 0. response times were analyzed with frequentist mixed-effects linear regression and accuracy was analyzed with mixed-effects logistic regression. in experiment 1, the critical fixed effects were state-of-affairs (sum-coded, –0.5 actual/0.5 alternate), polarity (sum-coded, –0.5 affirmative/0.5 negative) and their interaction. in experiment 2, the critical fixed effects were state-of-affairs (sum-coded, –0.5 actual/0.5 alternate), distance (sum-coded, –0.5 close/0.5 far) and their interaction. trial order was added as an additional (centered) fixed effect. pairwise comparisons were performed using the emmeans package (lenth 2017). the random structure of the models initially included intercepts and slopes for the critical fixed effects and their interaction. when a model failed to converge, its random effect structure was simplified following the recommendations in barr et al. (2013). for the linear models, p-values were computed using satterthwaite’s approximation for denominator degrees of freedom (kuznetsova, brockhoff & christensen 2013). 3. results. experiment 1 replicated the finding of a reduced sensitivity to truth-value in negated sentences: response times were faster for pictures showing actual vs. alternate states in affirmative, but not in (close distance) negative sentences, resulting in a significant state-of-affairs×polarity interaction (table 2 and figure 1). response times were also faster for pictures following affirmative vs. negative sentences. the accuracy analysis showed fewer errors for pictures of actual than alternate states. this effect was significant in affirmative and negative sentences, but it was numerically smaller in negative sentences, consistent with the reduced truth-value sensitivity in response times. proceedings of elm 3: 214-223, 2025 sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, and barbara kaup: insensitivity to truth-value in negated sentences: does linear distance matter?. 218 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ experiment 2 used the same close distance negated sentences as experiment 1, but the affirmative sentences were replaced with negated sentences in which the negator had a farther linear distance from the adjectival predicate. contrary to expectations, there was no evidence that sensitivity to truth-value was increased in the far distance negated sentences (i.e., non-significant state-of-affairs×distance interaction in response times). thus, we did not find that the distance between the negator and the predicate modulated sensitivity to truth-value in the response times of negative sentences. the response times only showed faster picture recognition times for close vs. far distance negated sentences. the accuracy analysis revealed fewer errors for pictures of actual vs. alternate states, but pairwise comparisons revealed that this effect was only significant in the close distance conditions—thus replicating the pattern seen with these sentences in experiment 1. response time accuracy β se t p β se z p experiment 1 intercept (grand mean) –1.28 0.04 –29.58 <.001 3.40 0.23 14.58 <.001 trial order –0.00 0.00 –14.40 <.001 0.03 0.00 8.94 <.001 state-of-affairs 0.05 0.02 2.66 .011 –1.98 0.37 –5.32 <.001 polarity 0.05 0.02 2.70 .010 0.22 0.23 0.98 .328 state-of-affairs×polarity –0.08 0.04 –2.40 .020 0.70 0.40 1.75 .080 soa: aff. sentences 0.09 0.03 3.24 .002 –2.33 0.43 –5.51 <.001 soa: neg. sentences 0.00 0.02 0.32 .746 –1.63 0.42 –3.85 <.001 experiment 2 intercept (grand mean) –1.22 0.05 –26.95 <.001 4.08 0.28 14.41 <.001 trial order –0.00 0.00 –15.33 <.001 0.02 0.00 5.57 <.001 state-of-affairs 0.00 0.02 –0.12 .907 –1.19 0.32 –3.78 <.001 distance 0.06 0.02 3.82 <.001 0.08 0.20 0.39 .694 state-of-affairs×distance –0.01 0.03 –0.49 .625 1.13 0.43 2.63 .009 soa: close distance 0.00 0.02 0.24 .809 –1.76 0.39 –4.54 <.001 soa: far distance –0.01 0.02 –0.42 .678 –0.23 0.38 –1.67 .094 table 2: results of the statistical analysis. abbreviations: aff. = affirmative, neg. = negative, soa = state-of-affairs. estimates are expressed in reciprocal milliseconds for response time and log odds for accuracy. proceedings of elm 3: 214-223, 2025 sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, and barbara kaup: insensitivity to truth-value in negated sentences: does linear distance matter?. 219 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: descriptive summary of the response times of correct responses (top row) and accuracy (bottom row), averaged across items and participants. error bars show 95% confidence intervals. abbreviations: neg. = negative. 4. discussion. we conducted two sentence-picture matching tasks to examine german speakers’ sensitivity to truth-value in negative sentences, as well as its modulation by the linear position of the negator. the findings of experiment 1 replicated the negation-by-truth-value interaction found in previous studies (kaup, lüdtke & zwaan 2005; tian, breheny & ferguson 2010). specifically, participants were faster judging pictures that truthfully represented the state-of-affairs described by the sentence, but this processing facilitation disappeared in negative sentences. this does not mean that participants were blind to truth-value: they showed fewer errors with actual than alternate pictures in both affirmative and negative sentences, which shows that truth-value affected proceedings of elm 3: 214-223, 2025 sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, and barbara kaup: insensitivity to truth-value in negated sentences: does linear distance matter?. 220 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ their answers. thus, the results of experiment 1 demonstrate that true sentences elicit more accurate responses—even in a task that does not require truth judgments—but that difficulties related to comprehending negative sentences can neutralize the effect of truth-value in processing time. experiment 2 focused on negative sentences and compared structures in which the negator appeared linearly close to the adjectival predicate vs. earlier in the sentence, i.e., separated by several words (and a clause boundary) from the predicate. in close distance negated sentences, we found fewer errors for actual than alternate pictures but no evidence of truth-value sensitivity in response times, thus replicating experiment 1. in long distance negative sentences there was no evidence of sensitivity to truth-value in either accuracy or response times. this fails to support the hypothesis that an early occurrence of the negator, which introduces more distance—and thus processing time—between the negator and the predicate, restores sensitivity to truth-value. with regard to 2-stage accounts of negation, our findings suggest that having more time to process the negator does not prevent the creation of a counterfactual interpretation when the adjectival predicate is read, or its suppression to proceed to the creation of an actual interpretation. previous findings indicated that the activation of a counterfactual interpretation depended on whether the linguistic and/or visual context made the actual and alternate interpretations similarly salient, or whether it introduced a question-under-discussion in which the truth of the affirmative counterpart was at issue (nieuwland & kuperberg 2008; orenes, beltrán & santamaría 2014; tian, breheny & ferguson 2010; tian, ferguson & breheny 2016; darley, kent & kazanina 2020). our study adds to previous research by demonstrating that giving participants more time to process the negator does not, by itself, reduce the activation of a counterfactual interpretation, at least when the target sentences are presented in isolation. our study conceptually replicates the event-related potential study of dudschig et al. (2019), and it demonstrates similar results using a different type of dependent measure and task (response times in a sentence-picture matching task) and a design in which participants’ decisions did not rely on detecting world knowledge violations. our study has some limitations, and it also leaves some open questions for future research. one limitation concerns the type of negation used in the far distance sentences, e.g., ‘it is not true that…’. while the close distance sentences simply negated a specific state of affairs, the far distance sentences introduced a type of metalinguistic negation that is typically used to reject a previous assertion (e.g., ‘the package is wrapped’). it is possible that this encouraged (rather than discouraged) the creation of a counterfactual affirmative interpretation and thus increased processing difficulty. this explanation would account for the finding that both long distance negative sentences elicited longer response times than the close distance sentences, consistent with higher processing effort. future research could address this possibility by using a different structure to manipulate the distance between the negator and the relevant predicate. an important open question concerns the potential relationship between the linear position of the negator and the pragmatic licensing of the sentence. specifically, it remains to be tested if the position of the negator would play a role if negative sentences were pragmatically licensed, e.g., if they had been presented in context, as opposed to in isolation. a second open question concerns the crosslinguistic generalizability of our findings. our target sentences were in german, a language in which sentential negation is located in a low fixed position (zeijlstra 2004), but in which the linear position of the negator “nicht” is variable (steube 2006; sudhoff 2008). for example, “nicht” follows finite verbs in main clauses, but precedes the verbal complex in subordinate clauses. moreover, definite determiner phrase objects and prepositional phrase adjuncts scramble across the negator, while other constituents do not (frey & pittner 1998). given proceedings of elm 3: 214-223, 2025 sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, and barbara kaup: insensitivity to truth-value in negated sentences: does linear distance matter?. 221 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the variable position of the negator, german comprehenders might adopt a conservative processing strategy and delay the interpretation of negation until the relevant predicate is encountered. thus, the linear position of the negator might not be a reliable cue in the comprehension of negation in languages like german, in which the base and linear position of negation differ. future research on languages in which the linear position of negation exhibits less variation (e.g., spanish, polish and basque) as well as languages with early occurring, preverbal negation such as spanish would be useful to address this possibility. 5. supplementary materials. data, analysis code and materials are publicly available at the open science framework: https://osf.io/x9ue3/. references barr, dale j. 2013. random effects structure for testing interactions in linear mixed-effects models. frontiers in psychology 4. https://doi.org/10.3389/fpsyg.2013.00328. box, g. e. p. & d. r. cox. 1964. an analysis of transformations. journal of the royal statistical society. series b (methodological) 26(2). 211–252. http://www.jstor.org/stable/2984418. carpenter, patricia a. & marcel a. just. 1975. sentence comprehension: a psycholinguistic processing model of verification. psychological review 82. 45–73. https://doi.org/10.1037/h0076248. darley, emily j., christopher kent & nina kazanina. 2020. a ‘no’ with a trace of ‘yes’: a mousetracking study of negative sentence processing. cognition 198. 104084. https://doi.org/10.1016/j.cognition.2019.104084. dudschig, carolin, ian grant mackenzie, claudia maienborn, barbara kaup & hartmut leuthold. 2019. negation and the n400: investigating temporal aspects of negation integration using semantic and world-knowledge violations. language, cognition and neuroscience. routledge 34(3). 309–319. https://doi.org/10.1080/23273798.2018.1535127. fischler, ira, paul a. bloom, donald g. childers, salim e. roucos & nathan w. perry jr. 1983. brain potentials related to stages of sentence verification. psychophysiology 20(4). 400– 409. https://doi.org/10.1111/j.1469-8986.1983.tb00920.x. frey, werner & karin pittner. 1998. zur positionierung der adverbiale im deutschen mittelfeld. linguistische berichte 176. 489–534. kaup, barbara & carolin dudschig. 2020. understanding negation: issues in the processing of negation. in viviane déprez & m. teresa espinal (eds.), the oxford handbook of negation, 635–655. oxford university press. https://doi.org/10.1093/oxfordhb/9780198830528.013.33. kaup, barbara, jana lüdtke & rolf a. zwaan. 2005. effects of negation, truth value, and delay on picture recognition after reading affirmative and negative sentences. proceedings of the annual meeting of the cognitive science society 27(27). https://escholarship.org/uc/item/19s068vb. kaup, barbara, jana lüdtke & rolf a. zwaan. 2006. processing negated sentences with contradictory predicates: is a door that is not open mentally closed? journal of pragmatics (special issue: processes and products of negation) 38(7). 1033–1050. https://doi.org/10.1016/j.pragma.2005.09.012. kuznetsova, alexandra, p. b. brockhoff & rune haubo bojesen christensen. 2013. lmertest: tests in linear mixed effects models. https://cran.r-project.org/package=lmertest. lenth, russell v. 2017. emmeans: estimated marginal means, aka least-squares means. https://cran.r-project.org/package=emmeans. proceedings of elm 3: 214-223, 2025 sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, and barbara kaup: insensitivity to truth-value in negated sentences: does linear distance matter?. 222 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ nieuwland, mante s. & gina r. kuperberg. 2008. when the truth is not too hard to handle: an event-related potential study on the pragmatics of negation. psychological science. sage publications inc 19(12). 1213–1218. https://doi.org/10.1111/j.14679280.2008.02226.x. oppenheimer, daniel m., tom meyvis & nicolas davidenko. 2009. instructional manipulation checks: detecting satisficing to increase statistical power. journal of experimental social psychology 45(4). 867–872. https://doi.org/10.1016/j.jesp.2009.03.009. orenes, isabel, david beltrán & carlos santamaría. 2014. how negation is understood: evidence from the visual world paradigm. journal of memory and language 74. 36–45. https://doi.org/10.1016/j.jml.2014.04.001. r development core team. 2024. r: a language and environment for statistical computing. vienna: r foundation for statistical computing. https://www.r-project.org. steube, anita. 2006. the influence of operators on the interpretation of dps and pps in german information structure. in valéria molnár & susanne winkler (eds.), the architecture of focus, 489–516. de gruyter mouton. https://doi.org/10.1515/9783110922011.489. sudhoff, stefan. 2008. zum relativen skopus von negation und fokuspartikeln im deutschen mittelfeld. in karin pittner (ed.), beiträge zu sprache und sprachen 6: vorträge der 16. jahrestagung der gesellschaft für sprache und sprachen, vol. 6, 317–328. münchen: lincom europa. tian, ye, richard breheny & heather j. ferguson. 2010. why we simulate negated information: a dynamic pragmatic account. quarterly journal of experimental psychology 63(12). 2305–2312. https://doi.org/10.1080/17470218.2010.525712. tian, ye, heather ferguson & richard breheny. 2016. processing negation without context – why and when we represent the positive argument. language, cognition and neuroscience. routledge 31(5). 683–698. https://doi.org/10.1080/23273798.2016.1140214. zehr, jeremy & florian schwarz. 2018. penncontroller for internet based experiments (ibex). https://doi.org/10.17605/osf.io/md832. zeijlstra, hedde. 2004. sentential negation and negative concord. utrecht: lot. proceedings of elm 3: 214-223, 2025 sol lago, petra schulz, esther rinke, elise oltrogge, carolin dudschig, and barbara kaup: insensitivity to truth-value in negated sentences: does linear distance matter?. 223 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ informational content vs. discourse orientation : experimental and computational perspectives grégoire winterstein, ghyslain cantin-savoie, samuel laperle, josiane van dorpe & nora villeneuve * abstract. the aim of this study is to investigate how human speakers and computational language models process (i) the informational content and (ii) the discourse orientation of natural language sentences. these two dimensions of meaning have received little attention outside theoretical literature, especially in the computational linguistics domain. to help fill this void, we present the results of four experiments that exploit the specific semantics of two french adverbs, namely presque (≃ ’almost’) and à peine (≃ ’barely’), which put these two dimensions of meaning at odds. each experiment focuses on one kind of population (humans or language models), and one kind of meaning (informational content or discourse orientation). our results show that humans are indeed sensitive to informational content and discourse direction, as assumed in the theoretical literature. language models exhibit a less transparent behavior. their performances in dealing with the semantics of presque appear in line with predictions based on the way these models are trained, but this does not extend to à peine. keywords. semantics ; pragmatics ; language models ; psycholinguistics ; entailment ; discourse semantics ; natural langue inference 1. introduction. a fair number of approaches to natural language semantics consider meaning in terms of informational content, understood as the part of an utterance’s meaning which conveys information about the world. informational content is typically understood in truth-conditional terms, and the meaning of a declarative sentence is often defined in such logical terms. however, natural meaning encompasses more than informational content, at least at a first glance. to see it, consider the (made-up) dialogue (1), taking place in a pub in which patrons need to get their beer at the bar. a finished their beer, gets up to get another one, and asks b: (1) a: can i bring you another beer? a. b: yes, i’m done with mine. b. b’: yes, i’m almost done with mine. c. b”: ? yes, i’m barely done with mine. several things are to be noticed in that example. first, (1-a) and (1-b) can be used to answer a’s question with the same effect, i.e., to indicate a desire for another drink. this is so, in spite of the fact that those two sentences are at odds in terms of informational content. this is because being *the authors wish to thank the members of the slic research group, as well as members of the audience of elm2 for their input and comments. this work was partially funded by the centre de recherche sur le langage, l’esprit et le cerveau (crlec-uqam). authors: université du québec à montréal, grégoire winterstein (winterstein.gregoire@uqam.ca) & ghyslain cantin-savoie, (cantin-savoie.ghyslain@courrier.uqam.ca) & samuel laperle, (laperle.samuel@courrier.uqam.ca) & josiane van dorpe, (vandorpe.josiane@courrier.uqam.ca) & nora villeneuve, (villeneuve.nora@courrier.uqam.ca). proceedings of elm 2: 299-309, 2023 c©2023 grégoire winterstein, ghyslain cantin-savoie, samuel laperle, josiane van dorpe and nora villeneuve published by the lsa with permission of the author(s) under a cc by license. 299 https://doi.org/10.3765/elm https://www.elm-conference.net/ almost done with one’s beer entails that one is not done with it (amaral 2007, jayez & tovena 2008). this observation could be accounted for by considering that even though being almost done entails not being done, it still conveys that one is proximally close to being done, which would be enough to warrant getting another round. however, things get more complicated when considering (1-c). here, the mention of being barely done does not support a positive answer to a’s question in (1), but rather will support a refusal. note that being barely done seemingly entails being done, admittedly not by a long stretch, but still in a way that is objectively more advanced than when one is almost done. thus, the explanation we sketched above, relying on a notion of proximality, cannot help us here: it seems that considerations of purely informational content cannot fully account for the discourse uses of the expressions at hand. such observations form the bedrock of ducrot and anscombre’s theory of argumentation within language (anscombre & ducrot 1983), which serves as inspiration for the present work. their main claim is that the meaning of an utterance has an argumentative component which determines the discourse orientation of that utterance, i.e. how a speaker can use that sentence in a discourse in order to support or refute other propositions. crucially, the discourse orientation of an utterance can be at odds with its informational content, and the examples in (1) illustrate such cases. this dissociation between the two kinds of content in those examples is directly imputable to the semantics of almost and barely. these two adverbs belong to the class of “argumentative operators”, i.e., natural language expressions whose semantics affect the argumentative orientation of their host utterance. the purpose of this research is to conduct an experimental investigation of this dichotomy in both the behaviour of human subjects and that of large language models, as used in contemporary applications of natural language processing. specifically, we want to test the sensitivity of both “populations” (humans and language models) to each type of meaning (informational content vs. discourse orientation) by testing their behavior on sentences that involve argumentative operators like almost and barely which put those two dimensions at odds. this was done through four distinct experiments involving the french adverbs presque and à peine, whose semantics is close to english almost and barely. each experiment targeted a specific population and type of meaning. experiments involving humans participants are reported in section 2, and those with language models in section 3. we conclude and discuss future directions for our work in section 4. 2. experiments: human participants. the experiments discussed in this section aimed at providing experimental support for the hypothesis that natural language meaning involves (at least) two distinct dimensions of meaning: (i) an “objective” informational content, which encodes descriptive and referential content, and (ii) a discourse orientation (or argumentative orientation) which determines how that sentence can be integrated in a larger discourse. in a way, those experiments are mostly a form of sanity check: there is already a large body of theoretical work which argues for such distinctions, and the differences seem intuitive enough. there also exists previous work with similar goals. in particular amaral (2007: chap. 3) discusses experimental work which also grounds the difference between the truth-theoretic entailments of proximal adverbs like almost and barely (which she calls the polar component of the adverbs), and the way those sentences can be used in a discourse (which she refers to as the proximal part of proceedings of elm 2: 299-309, 2023 grégoire winterstein, ghyslain cantin-savoie, samuel laperle, josiane van dorpe and nora villeneuve: informational content vs. discourse orientation. 300 https://doi.org/10.3765/elm https://www.elm-conference.net/ the meaning). amaral’s results are consistent with the ones we present below, though there are differences in the method used, which we highlight whenever relevant. in section 2.1, we describe the experiment that targets the sensitivity of participants to the informational content of utterances, and section 2.2 introduces the experiment about discourse orientation. we discuss the results of both experiments in section 2.3. 2.1. human participants and informational content (experiment 1). the goal of the first experiment was to determine whether participants were sensitive to the purported logical entailments of the french adverbs presque (≃ ’almost’) and à peine (≃ ’barely’). both adverbs modify gradable, therefore scalar, predicates, and so the entailments at hand are as follows: • an expression of the form presque x should indicate a degree of the relevant scale that is lower than the use of x alone • an expression of the form barely x should indicate a degree of the relevant scale that is higher than the use almost x, and equal or greater than the minimal degree at which x is taken to be true 2.1.1. method. the experiment was administered via an online questionnaire, hosted on the pcibexfarm platform (zehr & schwarz 2018). 43 participants were presented with a sentence in bold face and asked to place the eventuality described by the sentence along a scale with the use of a slider. the scale was presented below the sentence, and both extrema of the scale were spelt out. the scales were specific to each item, and designed so that the middle of the scale corresponds to the minimal degree such that the predicate used in the target sentence would be true. figure 1 illustrates an item. figure 1: item example from slider experiment (translations: top sentence (in bold): alex almost finished their beer. ; instructions: locate the situation described by the sentence above on the scale below ; left end of the slider : alex is 30 minutes away from finishing their glass. ; right end of the slider : alex finished their glass 30 minutes ago. on figure 1 the item to be judged is alex finished their beer. for that sentence, the scale is a temporal one, where the extrema are equally distant from the exact moment alex empties their glass: any place on the right hand part of the scale thus corresponds to an instant when alex has finished their beer. the experiment used 18 target items and 18 filler items. we considered three conditions across proceedings of elm 2: 299-309, 2023 grégoire winterstein, ghyslain cantin-savoie, samuel laperle, josiane van dorpe and nora villeneuve: informational content vs. discourse orientation. 301 https://doi.org/10.3765/elm https://www.elm-conference.net/ target items: • ∅: the target sentence with an unmodified predicate (as in figure 1) • presque: the target sentence with the predicate modified by presque • apeine: the target sentence with the predicate modified by à peine the presentation of items was pseudo-randomized using a latin-square design, so that every participant saw each condition 6 times, with differing orders and condition instantiations across participants. participants were recruited via snowball sampling, and each received a link that led them to one of two online questionnaires: either the one for this experiment, or the one for experiment 2 (described in section 2.3). participants were told they could only answer the questionnaire once, so that the sets of participants to the two experiments are disjoint. 2.1.2. results. the average z-scores (per participant) for each condition in experiment 1 are presented on figure 2. figure 2: average z-scores by condition in experiment 1 to measure the significance of the condition under study, we fitted linear mixed effect models with random intercepts for items and participants, and assessed the significance of our main factor via model comparison using likelihood ratio tests. we found a significant effect (χ2 = 34.741, p < 0.001), with the presque condition being scored significantly below the other two. the difference between apeine and ∅ was not significant. 2.2. human participants and discursive orientation (experiment 2). to get an account of the human ability to perceive the argumentative or discourse orientation of utterances, we ran an experiment in which we asked subjects to evaluate the naturality of a sentence in a given context. regarding our previously discussed hypothesis, our prediction is that the context in which proceedings of elm 2: 299-309, 2023 grégoire winterstein, ghyslain cantin-savoie, samuel laperle, josiane van dorpe and nora villeneuve: informational content vs. discourse orientation. 302 https://doi.org/10.3765/elm https://www.elm-conference.net/ the sentences with presque are judged to be natural are the same contexts in which bare sentences are judged natural, and that in such contexts the use of à peine will be judged odd. conversely, if we change the context in a way which allows à peine to be natural, we expect presque to sound degraded. 2.2.1. method. as for experiment 1, experiment 2 was administered via an online questionnaire, hosted on the pcibexfarm platform (zehr & schwarz 2018). 30 participants were presented with a short context, followed by a line break and the target sentence. participants were asked to rate the naturalness of the sentence in the given context, using a 7 point likert scale. in (2) we present a (translated) example of target item. (2) context : alex and jackie are enjoying a beer on a terrace of a bar. the waiter comes near them and asks : ”can i bring you another beer?”. alex answers : target : , i’m done with mine. as can be seen in (2), we manipulated two different factors. the first was the conclusion targeted by the speaker of the target sentence. in (2), the two possibilities were yes (pos conclusion) and no (neg conclusion). the other factor was similar to the one in experiment 1, i.e. the modification of the predicate in the target sentence (with three levels ∅, presque and apeine). this created 6 different conditions in total. the experiment included 18 target items and 36 distractors. the presentation of items was pseudo-randomized using a latin-square design, so that every participant saw each condition 3 times, with differing orders and condition instantiations across participants. participants were recruited in the same manner and at the same time as for experiment 1 (see section 2.1.1 above). 2.2.2. results. the results of experiment 2 are summarized in figure 3. model comparison between ordinal mixed models with random intercept and slopes for items and participants shows a significant effect of each of the factor under study, i.e. of the modification of the predicate (∅/presque/apeine, χ2 = 34.98, p < 1e−8) and the type of conclusion target (pos/neg, χ2 = 18.77, p < 1e−5). the interaction between the two factors is also significant (χ2 = 160.65, p < 1e−10). overall, the acceptability of context favoring positive conclusions was higher, but within each condition, preferences were reversed: within pos contexts, ∅ patterned with presque. in neg contexts, apeine was judged significantly more natural than the two other conditions, but some discrepancies were observed between ∅ and presque. 2.3. discussion. the results of experiments 1 and 2 largely support the predictions of theoretical models on the interpretation of proximal adverbs and are in line with previous experimental work in this domain. comparing our experiments with those of amaral (2007), we mostly differ in how we tested the sensitivity of participants to the entailments of the target sentences. while amaral asked participants for binary entailment judgment, our method relied on graded behaviors in experiment 1. this allowed us to observe that for certain items, the degree associated to ∅ was close to that of apeine while in other cases the degree of ∅ appeared higher than apeine’s. both configurations are compatible with a truth-conditional entailment pattern of à peine to its prejacent, but not necessarily one in terms of degrees, i.e. though à peine might be felt to entail the truth of its prejacent, it does not entail that the prototypical degree associated to the property holds proceedings of elm 2: 299-309, 2023 grégoire winterstein, ghyslain cantin-savoie, samuel laperle, josiane van dorpe and nora villeneuve: informational content vs. discourse orientation. 303 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 3: average z-scores of naturality judgments by condition in experiment 2 (which corresponds to the minimizing effect of à peine). overall, our results are thus consistent with the hypothesis that the linguistic knowledge of participants encompasses both the informational content of the utterances and their discourse orientation, and that they are able to tease those apart. 3. experiments: language models. computational language models are often taken to encode complex grammatical knowledge. in particular, contemporary models such as transformer based models (vaswani et al. 2017) (e.g. those of the bert family, devlin et al. 2019) are typically pre-trained on large amounts of data on relatively general tasks (such as predicting the nature of a masked token). this pre-training is supposed to capture general linguistic knowledge, which can later be fine-tuned for particular tasks. the pre-training of the models crucially relies on distributional information: the representations encoded by the models are rooted in the observation of co-occurrence patterns in the training data. in that sense, we expect the models to have captured information related to the discourse orientation of linguistic elements, since discourse orientation precisely relates to matters of cooccurrence at the discourse level. language models are however often fine-tuned for tasks like the natural language inference (nli) task in which the model is used to predict whether one sentence, called the premise, entails another, called the hypothesis (poliak 2020). such a task would thus rely on the manipulation of informational content rather than discourse orientation. experiments 1 and 2 are consistent with the hypothesis that human speakers are sensitive to, and can distinguish between, discourse orientation and informational content. we designed two additional experiments proceedings of elm 2: 299-309, 2023 grégoire winterstein, ghyslain cantin-savoie, samuel laperle, josiane van dorpe and nora villeneuve: informational content vs. discourse orientation. 304 https://doi.org/10.3765/elm https://www.elm-conference.net/ to test whether language models are also sensitive to both dimensions, using the same adverbs as in the experiments with human participants. we begin by describing how we built the dataset used in both experiments (section 3.1). we then introduce experiment 3 in which we test the predictions of fine-tuned language models on detecting the inferential patterns associated with each adverbs (section 3.2). in section 3.3, we present experiment 4 which targets the sensitivity of language models to discourse orientation. we discuss the results of these experiments in section 3.4. 3.1. dataset. our dataset consisted in naturally occurring sentences containing one of the two adverbs under study. the dataset was built by pseudo-randomly extracting sentences from a subset of the french version of wikipedia which contained either presque (1990 items) or à peine (2770 items). in addition to the target sentence, we also extracted the two preceding sentences to serve as context. every target sentence was then replicated in three conditions: • original: the original extracted sentence • bare: the sentence with the target adverb removed • switched: the sentence with the target adverb replaced by the other (i.e. presque replaced by à peine and vice-versa) every sentence was also tagged with the adverb that actually appears in it (irrespective of its original form), with the same possible values as in experiment 1 and 2, i.e. ∅/presque/apeine. 3.2. language models and informational content (experiment 3). to test the sensitivity of language models to informational content, we evaluated a model that was previously fine-tuned to the nli task on pairs of sentences from our dataset. if the model acquired some knowledge about the informational content of the adverbs under study, we expect it to predict that: • a sentence containing presque should contradict its prejacent • a sentence containing à peine should entail its prejacent 3.2.1. method. for the experiment, we used the camembert model for french (martin et al. 2020) which is pre-trained on french data and fine-tuned on the french part of the mnli dataset (williams et al. 2018). we tested the entailment patterns by taking original sentences as premises, and the bare sentences as hypothesis. for each premise/hypothesis pair the model gave us the probability that the hypothesis is true given the premise. 3.2.2. results. the results of the experiment are summarized in figure 4. as can be seen on the figure, the entailment patterns of the two adverbs starkly differ. for sentences containing presque the general prediction is that they entail their prejacent. on the other hand, sentences with à peine appear to be divided in two groups: those that entail their prejacent and those that do not, with few sentences in between. 3.3. language models and discursive orientation (experiment 4). to test the sensitivity of language models to discursive orientation we relied on pre-trained models, before any task-specific fine-tuning. the rationale of our choice is that the pre-training of language models aims at capturing matters related to the distribution of linguistic expressions, and that discursive proceedings of elm 2: 299-309, 2023 grégoire winterstein, ghyslain cantin-savoie, samuel laperle, josiane van dorpe and nora villeneuve: informational content vs. discourse orientation. 305 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 4: probability of entailment of the prejacent for sentences containing the target adverbs orientation boils down to a question of distribution. in experiment 2, we used the naturality of discourses to assess the sensitivity of human participants to discursive orientation. here, we will use the perplexity of a model towards the target sentences, a measure we define in the next subsection.. if the model acquired some information about the discursive orientation of the adverbs under study, we expect it to predict that: • replacing presque by à peine, or vice-versa should increase the perplexity • deleting presque should not significantly alter the perplexity, or do it in a way not as marked as a deletion of à peine 3.3.1. method. to test the sensitivity of language models to discourse orientation, we relied on the the french-gpt model (simoulin & crabbé 2021). the rationale for our choice was that, unlike bert models, gpt-like models are directional, and thus are suited to measuring the perplexity of the model about the continuation of a discourse. as mentioned above, we take the perplexity of the model as a proxy for the naturality of a discourse: the higher the perplexity, the less natural the discourse. we thus calculated the average perplexity of the model on our target sentences, in the three conditions original, switched and bare, using the two previous sentences as the context against which to measure the perplexity of the model. the perplexity measure uses the context before a particular word to assess its probability of occurrence. specifically, it is defined as “the exponentiated average negative log-likelihood of a sequence. if we have a tokenized sequence x = (x0, x1, . . . , xt), then the perplexity of xx is defined [as follows], where log pθ(xi|x0.15). figure 1. raw reading times (in ms). error bars show +/1 se in contrast, at spillover region 3, we see clear differences emerging. now, the people conditions are significantly slower than the you conditions (logrts: t=-2.85, p<0.005; residual 6 this difference could be perhaps taken as partial support for the perspective-taking hypothesis, but then – at least under the current formulation of the perspective-taking hypothesis – we should also find a reliable difference between you and people, which is not the case. furthermore, it’s worth noting that this difference is unlikely to be due to lexical frequency. people is less frequent than we and you (e.g. subtlex-us database, brysbaert & new 2009), and the large body of evidence showing that humans exhibit rapid sensitivity to lexical frequency indicates that more frequent words are easier to process than less frequent words (e.g. huizeling et al. 2022 for recent discussion) – which would lead us to expect the opposite rt pattern than what the residual logrt analyses show. proceedings of elm 2: 154-163, 2023 elsi kaiser and jesse storbeck: real-time processing of indexical and generic expressions. 160 https://doi.org/10.3765/elm https://www.elm-conference.net/ logrts: t=-2.708, p<.01) and also significantly slower than the we conditions (logrts: t=-2.84, p<.005; residual logrts: t=-2.73, p<.01). thus, these differences persist even when we control for word length differences by analyzing residual logrts. (here, |t| ≥ 2, so these effects survive even if we follow that criterion for assessing significance.) we suggest that it is unlikely that this slowdown at spillover region 3 in the people conditions is simply due to lexical frequency differences between the less-frequent people and the relatively more frequent forms you and we. this is because the preceding regions show no significant differences that could be plausibly construed as slowdowns caused by low lexical frequency. furthermore, at spillover region 3 we are already three words downstream from the critical word (region 0). especially when coupled with the patterns in the preceding regions and the large slowdown observed in region 3, this seems to argue against pure lexical frequency driving the slowdown at this point. rather, we suggest that the finding that sentences with people are read more slowly than sentences with you or we could be taken as preliminary evidence for the indexicality hypothesis. admittedly, more work is needed to further assess this idea. at spillover region 4, there is still a marginal slowdown in the people conditions relative to the you conditions in the logrt analyses (logrts: t=-1.692, p=0.09; residual logrts: t=-1.544, p=0.12) but no other differences. spillover region 5 shows no significant differences in either analysis. overall, then, at spillover region 3 we find that conditions with the subject people elicit longer reading times than conditions with we or you in subject position. 7. general discussion. this study is a preliminary foray into the real-time processing of nonanaphoric pronouns and generic expressions, and focuses on the use of you, we and people in covid-19 health messages. to the best of our knowledge, this is the first covid-related selfpaced reading study to test how different forms (you, we, people) impact reading time, which we take to reflect ease of processing. our results provide preliminary evidence for an increased processing load in public health messages with the non-indexical form people (relative to pronouns we and you) which we interpret as providing initial support for the indexicality hypothesis: the preliminary finding that health messages with people are processed more slowly than ones with we or you (even when we control for differences in word length) is compatible with the core idea of the indexicality hypothesis, which is that indexical linguistic expressions have some aspect of this special status ‘hard-wired’ into their meaning – so that it persists even in contexts where the forms are used generically/non-indexically – and can thus be processed more rapidly than expressions like people that are never indexical. to interpret people, some kind of additional representation needs to be evoked, and this representation does not ‘come for free’ as part of the speech situation, unlike the speaker and addressee referents of indexicals. the details of this line of reasoning still need to be worked out. it is worth noting that you and we mostly pattern alike in our data, suggesting that they are equally easy to process in the kinds of contexts we tested. in light of their shared indexical ‘potential,’ this may not be surprising – and indeed it is predicted by the indexicality hypothesis. open questions remain about how these results relate to experiments that did not collect processing data, including tu et al.’s (2021) findings about you being better at influencing people’s (hypothetical) behavior than we. more broadly, we emphasize that these issues, including the hypotheses we proposed in this paper, need further investigation but can be challenging to test due to the intrinsic lexical differences between the conditions. proceedings of elm 2: 154-163, 2023 elsi kaiser and jesse storbeck: real-time processing of indexical and generic expressions. 161 https://doi.org/10.3765/elm https://www.elm-conference.net/ also, it’s worth noting that, no matter what ultimately turns out to be the explanation for the differences in reading times, the fact that there do appear to be differences in reading times (which point to differences in ease of processing) suggests that this kind of research can have implications for the construction of easily-understood public health messages. one important future direction concerns investigating whether individual differences in people’s general attitudes about covid as well as their individual attitudes about specific kinds of mitigation behaviors (masking, vaccination, etc.) influence reading times. furthermore, independent of the covid-related questions, experimental work directly comparing the real-time processing of indexical vs. generic uses of you and we would be very informative; in the present study, we intentionally used a context that strongly favored generic interpretations and thus our results do not speak directly to how these forms are processed when they are unambiguously indexical. finally, in light of crosslinguistic variation in how different kinds of generic reference are expressed (see e.g. siewierska 2008, see also kaiser 2015 on finnish vs. english), crosslinguistic experimental work would also be very welcome. as a whole, the results reported in this paper can help further our understanding of how different non-anaphoric pronominal expressions – often neglected in prior models of pronominal processing – are processed in real time, and this kind of research potentially has practical implications for the construction of easily-understood public health messages. references baayen, r. harald, davidson, doug j., & bates, douglas. 2008. mixed-effects modeling with crossed random effects for subjects and items. journal of memory and language 59(4). 390412. bates, douglas, maechler, martin, bolker, ben, & walker, steve. 2015. fitting linear mixedeffects models using lme4. journal of statistical software 67(1). 1-48. braun, david. (2001). indexicals. in edward n. zalta (ed.), the stanford encyclopedia of philosophy. available at: http://plato.stanford.edu/entries/indexicals brunyé, tad t., ditman, tali, mahoney, carolone, augustyn, jason & taylor, holly. 2009. when you and i share perspectives. pronouns modulate perspective taking during narrative comprehension. psychological science 20(1). 27-32. brysbaert, marc & new, boris. 2009. moving beyond kucera and francis: a critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for american english. behavior research methods 41 (4). 977-990. ferguson, heather, apperly, ian & cane, james. 2017. eye tracking reveals the cost of switching between self and other perspectives in a visual perspective-taking task. the quarterly journal of experimental psychology 70(8). 1646-1660. hinterwimmer, stefan & brocher, andreas. 2018. an experimental investigation of the binding options of demonstrative pronouns in german. glossa: a journal of general linguistics 3(1): 77. 1-25. holmberg, anders, & phimsawat, on-usa. 2017. truly minimal pronouns. diadorim 19. 11-36. huizeling, eleanor, arana, sophie, hagoort, peter, & schoffelen, jan-mathijs. 2022. lexical frequency and sentence context influence the brain’s response to single words. neurobiology of language 3(1). 149-179. kaiser, elsi. 2015. impersonal and generic reference: a cross-linguistic look at finnish and english narratives. eesti ja soome-ugri keeleteaduse ajakiri. journal of estonian and finnougric linguistics, 6(2), 9-42. proceedings of elm 2: 154-163, 2023 elsi kaiser and jesse storbeck: real-time processing of indexical and generic expressions. 162 https://doi.org/10.3765/elm https://www.elm-conference.net/ kaiser, elsi. 2021. effects of linguistic manipulations on the comprehension of covid-19 health messages. talk at architectures and mechanisms for language processing (amlap) 2021, paris. kamio, akio. 2001. english generic we, you, and they: an analysis in terms of territory of information. journal of pragmatics 33(7). 1111-1124. keysar boaz, barr dale j., balin, jennifer a. & brauner jason s. 2001. taking perspective in conversation: the role of mutual knowledge in comprehension. psychological science 11(1). 32-38. kuznetsova, alexandra, brockhoff, per b., & christensen, rune h.b. 2017. lmertest package: tests in linear mixed effects models. journal of statistical software 82(13). 1-26. luke, steven g. 2017. evaluating significance in linear mixed-effects models in r. behavior research methods. 49. 1494-1502. moltmann, friederike. 2006. generic one, arbitrary pro, and the first person. natural language semantics 14. 257-281. orvell, ariana, kross, ethan & gelman, susan. 2020. “you” speaks to me: effects of genericyou in creating resonance between people and ideas. proceedings of the national academy of sciences 117(49). 31038-31045. siewierska, anna. 2008. ways of impersonalizing: pronominal vs. verbal strategies. in maría de los angele gómez-gonzález, lachlan mackenzie, and elsa m. gonzáles álvarez (eds.), current trends in contrastive linguistics: functional and cognitive perspectives (studies in functional and structural linguistics), 3-26. amsterdam: john benjamins. tu, ke c., chen shirley s. & mesler rhiannon, m. 2021. “we” are in this pandemic, but “you” can get through this: the effects of pronouns on likelihood to stay-at-home during covid-19. journal of language and social psychology 40(5-6). 574-588. warren, tessa & gibson, edward. 2002. the influence of referential processing on sentence complexity. cognition 85(1). 79-112. warren, tessa & gibson, edward. 2005. effects of np type in reading cleft sentences in english. language and cognitive processes 20(6). 751-767. whitley, stanley. 1978. person and number in the use of we, you and they. american speech 53(1). 18-39. zehr, jeremy, & schwarz, florian. 2018. penncontroller for internet based experiments (ibex). https:// doi.org/10.17605/osf.io/md832 proceedings of elm 2: 154-163, 2023 elsi kaiser and jesse storbeck: real-time processing of indexical and generic expressions. 163 https://doi.org/10.3765/elm https://www.elm-conference.net/ tracking the activation of scalar alternatives with semantic priming eszter ronai & ming xiang* abstract. from an utterance of mary ate some of the deep dish, hearers frequently infer that mary didn’t eat all of the deep dish. similarly, an utterance of the movie is good might lead hearers to conclude that the movie isn’t excellent. these inferences are instances of scalar implicature (si). the standard assumption is that si arises via hearers’ reasoning about alternative utterances that the speaker could have said, but did not. in particular, hearers are taken to consider stronger alternatives such as all (or mary ate all of the deep dish) and excellent (or the movie is excellent) and derive their negation. in this study, we investigate the psycholinguistic reflexes of this inferential process. we use semantic priming with lexical decision to test whether lexical alternatives such as all and excellent are retrieved and activated in the processing of si-triggering sentences. the results of our experiments indeed suggest that alternatives play a role in the processing of si, though a number of empirical puzzles remain. keywords. pragmatics; scalar implicature; alternatives; semantic priming 1. introduction. scalar implicature (si), exemplified in (1), is a classic example of utterances receiving an enriched meaning that goes beyond their literal meaning. (1) mary ate some of the deep dish. literal meaning: mary ate at least some of the deep dish. si-enriched meaning: mary ate some, but not all, of the deep dish. it is commonly assumed that the inferential process that gives rise to si involves hearers reasoning about informationally stronger unsaid alternatives. for example, upon encountering the utterance in (1), hearers consider the stronger statement mary ate all of the deep dish, which the speaker could have said, but did not. hearers can infer as an si the negation of this stronger alternative, enriching the literal meaning of (1) with mary didn’t eat all of the deep dish. this process can be viewed as an interaction of the quality and quantity maxims (grice 1967). it is an open question what psycholinguistic mechanisms underlie the inferential process that gives rise to si. to address this, this paper uses semantic priming with lexical decision to test whether unsaid alternatives are retrieved and activated in the processing of si-triggering sentences. the general logic of our experiments is to probe whether alternatives like all are recognized with facilitated reaction times in a lexical decision task when they follow a relevant si-triggering sentence like (1). our findings suggest that comprehenders indeed activate the alternatives that theories of si take them to reason about; in other words, lexical scales are psychologically real. *for helpful discussion and feedback, we would like to thank ira noveck and nicole gotzner, as well as audiences at elm 2 and at the laboratoire de linguistique formelle linglunch. this material is based upon work supported by the national science foundation under grant no. #bcs-2041312. all mistakes and shortcomings are our own. stimuli, data, and the scripts used for data visualization and analysis can be found in the following osf repository: https://osf.io/wga25/?view_only=cefc447bc6e649aeb4815de958d71597 authors: eszter ronai, the university of chicago (ronai@uchicago.edu) & ming xiang, the university of chicago (mxiang@uchicago.edu). proceedings of elm 2: 229-240, 2023 c©2023 eszter ronai and ming xiang published by the lsa with permission of the author(s) under a cc by license. 229 https://doi.org/10.3765/elm https://www.elm-conference.net/ though much research has concentrated on the scale and the corresponding some but not all si, other lexical items also form scales. as (2) demonstrates, an utterance containing good can invite reasoning about the stronger alternative excellent, and lead to a good but not excellent si. recently, attention has turned to investigating a wider range of scales, with findings uncovering variation in the likelihood of si: e.g., the si in (2) is much less likely to arise than the one in (1) (i.a. van tiel et al. 2016). in our investigation of alternative activation, we capitalize on this phenomenon of scalar diversity, and our priming experiments will test 60 different scales. (2) the movie is good. literal meaning: the movie is at least good. si-enriched meaning: the movie is good, but not excellent. this paper is structured as follows. section 2 reviews previous work on the processing of alternatives. we then present four semantic priming experiments: experiment 1 is a replication unrelated to si (section 3); experiment 2 tests scalar alternatives without sentential context (section 4); experiment 3 tests alternatives in the context of si-triggering sentences (section 5); finally, experiment 4 tests focus alternatives (section 6). in section 7, we discuss what relevance priming results can have for different theories of si, as well as some remaining empirical puzzles. 2. alternatives in language processing. alternatives are pervasive in (the modeling of) semanticpragmatic meaning. correspondingly, they have generated a lot of interest in psycholinguistics, with various experimental paradigms being used to probe what kind of mental representations they have. in this brief section, we will concentrate on focus and scalar alternatives. for comprehensive overviews of alternative processing, including also alternatives involved in negation and counterfactuals, see repp & spalek (2021) and gotzner & romoli (2022) (and references therein). sentential focus marks new or emphasized information in a sentence. this information is often provided in (implicit) contrast to possible other alternatives (rooth 1992, 1985). in english, focus can be marked for instance by placing a prominent accent on a word: (3) mary ate deep dish. (3) conveys not only that mary ate deep dish, but also that mary did not eat anything else from among a set of contextually determined of alternatives, e.g., {lasagne, salad, ...}. in successful comprehension, hearers infer this set of contrastive alternatives as intended by the speaker. a growing number of studies have found that focus alternatives such as lasagne above are activated in the processing of sentences like (3). our experiments testing scalar alternatives are modeled after studies that used semantic priming to test the processing of sentential focus and the activation of focus alternatives. in particular, husband & ferreira (2015) (following braun & tagliapietra 2009; see also gotzner et al. 2016, yan & calhoun 2019) used lexical decision with cross-modal priming. in this study, participants listened to sentences such as (4), and had to make a decision about whether a visually presented target word was a word of english. (4) the murderer killed the nurse last tuesday night. the prime in each sentence was the focused element (nurse in (4)), while the targets in the lexical decision task were contrastive semantic associates (focus alternatives, e.g. doctor), non-contrastive semantic associates (e.g. clinic) and unrelated words. the study found early activation (i.e. faciliproceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 230 https://doi.org/10.3765/elm https://www.elm-conference.net/ tated lexical decision reaction times) of both contrastive and non-contrastive semantic associates in sentences where nurse was focused. importantly, however, later activation (after a longer stimulus onset asynchrony) was only found for contrastive alternatives. this suggests that comprehenders establish the proper set of focus alternatives during comprehension, which then allows them to draw the relevant inferences intended by the speaker, i.e., that the murderer did not kill the doctor. though previous priming studies are the most relevant for our paper, evidence of the activation of focus alternatives comes from a much larger body of work. existing studies have tested many different ways of marking sentential focus, including not just intonation, but focus particles (e.g., only, also), cleft sentences, or font emphasis. they have also successfully used a variety of experimental paradigms, such as probe recognition, delayed recall, change detection, or visual world eye-tracking. the readers are referred to i.a., sanford et al. (2009), kim et al. (2015), fraundorf et al. (2010), fraundorf et al. (2013), spalek et al. (2014), and gotzner & spalek (2017). priming has also been used to investigate scalar alternatives. de carvalho et al. (2016) used lexical decision with subliminal priming to see if one member of a scale (e.g., some) activates the other (all). participants were visually presented with a prime word for 34ms and then had to decide whether the following visually presented target was a word of english. the authors’ goal was to adjudicate between different theories of si. they made the assumption that under a neo-gricean account of si, which relies on lexically given horn-scales, the stronger alternative all is always needed in the processing of the weaker term some, but not vice versa. this makes the prediction that some would prime all more than all primes some. a post-gricean account such as relevance theory, on the other hand, does not assign special significance to lexical scales. the authors therefore predicted that under post-gricean accounts, any priming effect should merely be due to semantic relatedness and not show asymmetry, i.e., some and all would prime each other equally. the findings are in line with the neo-gricean account. an important difference between this study and ours (as well as the literature on focus alternatives), is that de carvalho et al. tested whether scalar terms prime each other in the absence of any sentential context. in contrast, what we are primarily interested in is whether scalar alternatives are primed in si-triggering sentences. another relevant priming study is by schwarz et al. (2016), whose research question addressed not whether scalar alternatives are activated in the processing of si. rather, they tested the hypothesis that by presenting an alternative as the prime, its salience is increased, which might lead to more likely and faster si calculation. ultimately, the findings did not support this hypothesis. lastly, there is also a growing body of work that uses not lexical, but structural priming to investigate si, see i.a., rees & bott (2018), bott & frisson (2022), and bott & chemla (2016). 3. experiment 1: replication of thomas, et al. (2012). given that priming experiments are typically conducted in person in a lab setting, we first carried out a replication study of the basic semantic priming effect in thomas et al. (2012), in order to validate our web-based methodology. 3.1. participants and task. 50 native speakers of american english participated in an online experiment on the pcibex platform (zehr & schwarz 2018). participants were recruited on prolific and compensated $2. native speaker status was established via a language background questionnaire; payment was not conditioned on the participant’s response. participants were removed if their lexical decision accuracy was below 90%. data from 39 participants is reported below. experiment 1 was a semantic priming with lexical decision experiment. participants had to proceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 231 https://doi.org/10.3765/elm https://www.elm-conference.net/ decide whether a word they saw was a word of english or not; this word was the target. they had to press the f key for “non-word” and the j key for “word”. the primary dependent variable of interest is their reaction time in making this lexical decision. participants were instructed to make a decision as fast as possible, while remaining accurate. before making a lexical decision on the target, participants saw the prime word. there were two within-participants conditions: in the “related” condition, the target word (e.g., boy) was preceded by a prime word that was semantically related to it (girl). in the “unrelated” condition, the same target (boy) was preceded by a prime that was not semantically related (boulevard). primes appeared in uppercase and targets in lowercase. participants first saw a fixation cross displayed for 350ms. it was then followed by 400ms of a blank screen. after that, the prime word appeared for 150ms. the presentation of the prime word was followed by another 650ms blank screen —this is the stimulus onset asynchrony (soa), i.e., the time between the offset of the prime and the onset of the target. finally, participants saw the target word, which they had to make the lexical decision on. if a participant did not make a lexical decision within 3000ms of the onset of the target, the experiment moved on to the next trial. the related condition in experiment 1 used 60 prime-target pairs from the “symmetrical associates” in thomas et al. (2012; p. 640, table a1). these pairs are symmetrical because the prime has a meaning that evokes the target, and vice versa, e.g., girl-boy, circle-square, salt-pepper. the unrelated prime words were randomly selected from the “forward associates” in thomas et al. (2012; p. 640, table a1). the experiment also included 60 fillers items with non-word targets. of the filler targets, 30 were 4-10/11 letter pseudohomophones that we generated from the arc nonword database (rastle et al. 2002) —e.g., spraized, knewed —, and 30 were non-words from lupker & pexman (2010; p. 1282, standards-1) —e.g., cleam, dronk. the experiment started with 10 practice items: 5 words and 5 non-words. for the first 4 practice items only, participants saw reminder labels that the f key corresponded to “non-word” and j to “word”. 3.2. hypothesis and predictions. we predict to replicate thomas et al.’s result: shorter lexical decision reaction times (rt) in the related, as compared to the unrelated condition. in the related condition, the target has been preceded by a semantically similar word, which should activate its meaning and facilitate its recognition. in the unrelated condition, the prime would not activate the target, which is then recognized at a “baseline” speed, related to its frequency, length, etc. (see also i.a., swinney 1979, swinney et al. 1979 for classic findings of semantic priming.) 3.3. results and discussion. data points with incorrectly answered lexical decision responses (i.e., a “non-word” response) were excluded, removing 2.09% of the data. figure 1 shows mean rt (and standard error) by condition. for the statistical analysis, a linear mixed effects regression model was fit (lme4 package in r; bates et al. 2015), predicting rt on the target word by condition (“related” vs. “unrelated”). the fixed effects predictor condition was sum-coded (related: -0.5 and unrelated: 0.5). random intercepts and slopes were included for participants and items. rts in the related condition were found to be significantly shorter than in the unrelated condition (estimate=25.51, se=8.65, t=2.95, p<0.01). that is, participants recognized words faster when they have been primed by a semantically related word. this successfully replicates thomas et al.’s in-lab results, validating the web-based setup —though we must note that the magnitude of the priming effect (i.e., the difference in rt between the related and unrelated conditions) was smaller in our experiment. proceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 232 https://doi.org/10.3765/elm https://www.elm-conference.net/ 586.9 613.9 580 590 600 610 620 condition r t ( m s) figure 1: results of experiment 1: replication of thomas et al. (2012) 4. experiment 2: lexical priming. experiment 2 was also a lexical semantic priming experiment, but this time the primes were weaker scalar terms from a scale (e.g., some, good), while the targets were stronger alternatives (e.g., all, excellent). importantly, scalar term primes were not placed in a sentential context, where si could have been calculated. this means that if we see semantic priming in experiment 2, that will reflect semantic similarity, not si calculation. experiment 2 therefore provides a baseline for later experiments that test inference-triggering sentences. 4.1. participants and task. 49 native speakers participated for $2 compensation. recruitment and screening was identical to experiment 1, including the exclusion criterion. data from 44 participants is reported below. capitalizing on scalar diversity, experiment 2 used 60 different lexical scales as critical items. each item consisted of a pair of scalar terms where the stronger term asymmetrically entails the weaker one —see ronai & xiang (to appear) for how this scale set was constructed. prime words in the “related” condition were weaker scalar terms like good, while targets were stronger alternatives like excellent. in the “unrelated” condition, the primes were instead unrelated words such as foreign. unrelated primes were generated to satisfy two criteria. first, they had to fit into the sentence frames used in experiments 3-4: given the sentence the movie is good, foreign was chosen, since the movie is foreign is also an acceptable sentence. second, unrelated primes had to have sufficiently low semantic similarity with the target (average cos(θ)=0.138). this was operationalized using vector semantics, specifically the glove model and spacy word embeddings. other than the critical test items, experiment 2 was identical to experiment 1 in its task, procedure (including timing parameters such as soa), filler and practice items. 4.2. hypothesis and predictions. if pairs of scalar terms (e.g., good-excellent) are semantically similar enough to lead to priming, then the results of experiment 2 should pattern similarly to experiment 1 —we should see shorter rts in the related condition than in the unrelated condition. 4.3. results and discussion. data points with incorrectly answered lexical decision responses (“non-word”) were excluded, removing 2.35% of the data. figure 2 shows mean rt (and standard error) by condition. statistical analysis was identical to experiment 1, except for the random effects structure, which included random intercepts for participants and items and random slopes for participants. the statistical analysis revealed no significant difference between rts in proceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 233 https://doi.org/10.3765/elm https://www.elm-conference.net/ the related and unrelated conditions (estimate=11.46, se=9.94, t=1.15, p=0.26). 678.9 685.6 670 675 680 685 690 695 related unrelated condition r t ( m s) figure 2: results of experiment 2: lexical priming experiment testing scalar alternatives that is, targets in the related condition were not recognized significantly faster than in the unrelated condition. this suggests that pairs of scalar terms do not lead to semantic priming when the words are presented in isolation, in the absence of any sentential context. therefore, we will be able to conclude that any priming effect we find in sentential experiments (experiments 3-4) is due to inference processing and alternative retrieval, not just mere meaning similarity. 5. experiment 3: sentential priming. having seen in experiment 2 that weaker scalar terms do not prime stronger alternatives in isolation, experiment 3 turns to priming effects in sentential contexts. here, we test whether stronger scalar alternatives are retrieved and activated in the processing of sentences that lead to si calculation, and are taken to involve reasoning about alternatives. 5.1. participants and task. 50 native speakers participated for $3.20/3.50 compensation. recruitment and screening was identical to experiment 1. data from 46 participants is reported. experiment 3 was also a lexical decision task with two within-participants conditions (related vs. unrelated). target words were the same scalar terms as in experiment 2. importantly, however, primes were now full sentences: the prime words from experiment 2 were embedded in a sentential context. that is, while for the scale experiment 2 used the word good as a prime, in experiment 3 good appeared in a sentence: the movie is good. similarly, in the unrelated condition, the unrelated words were embedded in a sentential context, e.g., the movie is foreign. each trial started with a 350ms fixation cross. this was followed by 400ms of a blank screen. after that, prime sentences were presented word-by-word, with each word being displayed for 350ms. there was a 650ms soa between the offset of the final word in the sentence (good/foreign) and the onset of the target (excellent). as before, if a lexical decision was not made within 3000ms of the onset of the target, the experiment moved on to the next trial. filler and practice targets used the materials of experiments 1-2, but the primes were sentences, not single words. 5.2. hypothesis and predictions. if stronger scalar alternatives like excellent are reasoned about and retrieved during si calculation, then we should see shorter rts in the related condition than in the unrelated condition. that is, excellent should be recognized faster following an sitriggering sentence where it serves as a stronger alternative. on the contrary, if lexical alternatives do not play a role in si processing, then there should be no rt difference between the conditions. proceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 234 https://doi.org/10.3765/elm https://www.elm-conference.net/ 5.3. results and discussion. excluding incorrectly answered lexical decision responses (“non-word”) removed 1.7% of the data. figure 3 shows mean rt (and standard error) by condition. statistical analysis was identical to experiment 2. rts in the related condition were found to be significantly shorter than in the unrelated condition (estimate=21.62, se=8.65, t=2.5, p<0.05); targets were recognized faster following an si-triggering sentence. 639.9 664.2 640 650 660 670 related unrelated condition r t ( m s) figure 3: results of experiment 3: sentential priming experiment testing scalar alternatives experiment 3’s findings therefore show that a stronger scalar alternative like excellent is recognized faster as a word of english when it has been preceded by a sentence like the movie is good, which can trigger the not excellent si. this, in turn, suggests that in the processing of such an si-triggering sentence, comprehenders retrieved and activated the relevant stronger scalar alternative. in the unrelated condition, on the other hand, such alternative targets would not have been activated in the processing of the prime sentence, and were therefore recognized with a baseline rt. let us recall that these findings cannot receive an explanation simply in terms of semantic similarity. the prime sentences were identical across the related and unrelated conditions up until the critical word (the movie is x). and as for the critical word (the weaker scalar term good vs. the unrelated word foreign), experiment 2 demonstrated that their difference in meaning, and the similarity between good and excellent does not, in itself, lead to semantic priming. there is, however, one important caveat: we cannot be certain that what the priming effect is evidence for is the retrieval of specific lexical items (e.g., excellent). it is also possible that the observed facilitation in rts is due to a more general activation of semantic features associated with the stronger alternative state. for instance, given the sentence the movie is good, participants might have considered the stronger alternative state where the movie is more than good, but without necessarily reasoning about the specific alternative excellent —this might still result in the observed effect. we also cannot be certain that participants in experiment 3 actually calculated the sis (e.g., not excellent), since the experiment did not include a task to probe si calculation. 6. experiment 4: sentential priming with only. numerous studies have shown that focus alternatives are activated in sentence processing (section 2). as another baseline to experiment 3, we therefore conducted an experiment where prime sentences also included the focus particle only. 6.1. participants and task. 50 native speakers participated for $3.20 compensation. recruitment and screening was identical to experiment 1. data from 43 participants is reported below. experiment 4 was identical to experiment 3 in its task and procedure (including timing paramproceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 235 https://doi.org/10.3765/elm https://www.elm-conference.net/ eters), with one difference: critical items were modified such that prime sentences in the related condition also included the focus particle only. that is, before participants made a lexical decision on a stronger alternative target such as excellent, they saw the prime sentence the movie is only good (presented word-by-word). the unrelated conditions, as well as filler and practice items were identical to experiment 3, i.e., they were not modified to include the word only. 6.2. hypothesis and predictions. the exclusion of focus alternatives is encoded semantically (rooth 1985, 1992), while alternatives in si are excluded pragmatically, in a cancellable way. given that experiment 3 already revealed evidence for alternative activation in si, we can predict to see similar effects in experiment 4. moreover, this is also what we expect based on previous work that has tested focus alternatives in a variety of experimental paradigms. this leads to the strong prediction that rts should be shorter in the related condition than the unrelated condition. 6.3. results and discussion. excluding incorrectly answered lexical decision responses (“non-word”) removed 1.98% of the data. statistical analysis was identical to experiment 2. figure 4 shows mean rt (and standard error) by condition. rts in the related condition were significantly shorter than in the unrelated condition (estimate=24.47, se=8.01, t=3.06, p<0.01). 628 654.3 620 630 640 650 660 related unrelated condition r t ( m s) figure 4: results of experiment 4: sentential priming experiment testing focus alternatives like experiment 3, experiment 4 revealed facilitation for stronger alternative targets in the related condition. that is, the prime sentence the movie is only good led to a faster recognition of the word excellent. semantic theory holds that sentences including focus (signalled e.g., by only) encode the exclusion of alternatives —our findings provide evidence that excluded alternatives have a processing correlate. this is in line with existing work reviewed in section 2. 7. summary of findings. the experiments in this paper (especially experiment 3) show evidence that stronger scalar alternatives (all, excellent) are retrieved and activated in the real-time processing of si. this informs our understanding of the mental representations behind pragmatic reasoning. classic gricean accounts of si hold that comprehenders reason about, and derive the negation of, relevant informationally stronger alternatives that the speaker could have said, but did not say. our findings suggest that this reasoning process also has processing correlates: relevant alternatives are activated when hearers process si-triggering sentences. in what follows, we discuss to what extent our findings might further inform theory, as well as some remaining puzzles. proceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 236 https://doi.org/10.3765/elm https://www.elm-conference.net/ 7.1. relevance for theory?. previous work on the activation of scalar alternatives has used findings to adjudicate between neoand post-gricean accounts (de carvalho et al. 2016). here, we briefly review these accounts and discuss what relevance priming evidence might have for them. neo-gricean accounts typically assume that hearers infer the negation of informationally stronger alternatives that the speaker could have said —e.g., because form a lexical scale, and all is stronger than some, hearers derive not all upon encountering some. these alternatives are determined via the lexicon (i.a., horn 1972) or the grammar (i.a., katzir 2007). on the other hand, post-gricean accounts such as relevance theory (i.a., sperber & wilson 1995) do not attach special importance to lexical scales. instead, the final interpretation of an utterance is obtained through a context-based enrichment process or ad hoc concept construction (in the case of si, strengthening). one seemingly straightforward way to interpret the findings of our experiments, then, is that they support neo-gricean accounts of si, since those take hearers to reason about particular lexical alternatives. we could also say that our results are not predicted by theoretical accounts of si that dispense with lexical scales, such as relevance theory. however, we believe that at least two issues arise with this interpretation of the results, having to do with what predictions different theories of si may make for priming data. first, it is not clear whether neo-gricean accounts would predict that stronger alternatives are activated when the weaker scalar term is presented in isolation. one the one hand, we could assume that lexical scales should only be relevant in language processing when alternatives are actually reasoned about, in the context of an si-triggering utterance. if so, then the findings of this paper do indeed support neo-gricean accounts, since we found priming in a sentential context (experiment 3), but not in isolation (experiment 2). on the other hand, if lexical scales are hardwired into the lexicon, then we might predict that pairs of scalar terms prime each other even in the absence of an si-triggering sentence. as mentioned, de carvalho et al. (2016) made a prediction along these lines: that the weaker scalar term primes the stronger alternative asymmetrically. following this reasoning, our findings in fact do not fully support neo-gricean accounts, since no priming was found when the scalar terms occurred in isolation (experiment 2). the finding that alternatives are only activated in (a sentential) context (experiment 3) could even be argued to support relevance theory, where si calculation occurs only when there is sufficient support from context. second, the activation of alternatives may signal either that alternatives were retrieved for the si calculation process to occur, or we might see activation as a by-product of the si calculation process. broadly speaking, neo-gricean accounts would assume that for si to arise, particular lexical items from lexical scales are reasoned about —this would predict the priming effect we found in experiment 3. but it is also possible that the priming effect we see is epiphenomenal. on post-gricean theories, hearers still calculate si, even though lexical scales do not play a special role. and once hearers have reached the si-enriched meaning (≈the movie is no more than good.), this could then lead to the observed priming effect, even if the stronger alternative excellent was not retrieved in the first place. these two ways of interpreting the priming findings are related to the issue discussed in section 5.3: whether facilitated rts constitute evidence that a specific lexical item excellent was retrieved, or whether they simply suggest that some semantic features related to the alternative state “the movie is more than good” were activated. given the above, we would argue that as things stand, no firm conclusions can be reached proceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 237 https://doi.org/10.3765/elm https://www.elm-conference.net/ about the validity of neo-vs. post-gricean accounts based on priming evidence. 7.2. remaining empirical puzzles. some open questions remain. first, we saw that experiments 3 and 4 pattern alike: rts were facilitated in the related condition. an additional statistical analysis on the combined experiment 3-4 data set also found no effect of experiment (estimate=9.51, se=22.53, t=0.42, p=0.67). this suggests that alternatives like excellent are similarly activated no matter whether the sentence that is processed is the movie is good or the movie is only good. this is despite the fact that si excludes alternatives pragmatically, but in the case of focus, alternative exclusion is semantic. indeed, the movie is only good is significantly more likely to lead to the not excellent inference than the movie is good (ronai & xiang 2022). the lack of a difference between the current experiment 3 and 4 suggests that the activation of alternatives, as measured via priming, does not track the rate of inference from the corresponding sentences: more robust inference calculation does not correspond to stronger priming. second, experiment 2 served to rule out the possibility that a priming effect in experiment 3 would reflect mere meaning similarity, rather than the processing correlate of reasoning about alternatives. for this reason, it is a welcome result that experiment 2 revealed no effect. but it is itself a puzzle why we found semantic priming in experiment 1 but not in 2. differences in (vector) semantic similarity cannot provide an explanation: table 1 shows that average prime-target similarity (based on glove/spacy) in the related vs. unrelated condition is quite similar in the two experiments. it is perhaps possible that the nature of the relationship between prime and target is different in thomas et al. (2012) (and classic semantic priming studies) than in the case of scalar items: there is an intuitive sense in which girl-boy and salt-pepper are closer semantically than good-excellent and some-all. future research should probe the exact source of semantic priming. cosine similarity related condition unrelated condition experiment 1 (replication of thomas et al.) 0.605 0.126 experiment 2 (scalar items in isolation) 0.707 0.138 table 1: average semantic similarity between prime-target pairs in experiments 1 and 2. lastly, as mentioned in the introduction, likelihood of si varies across lexical scales. this, however, does not correspond to a systematic difference in priming. to test this, we looked at the correlation between the experiment 3 priming effect across items and the corresponding si calculation rates (from ronai & xiang to appear, experiment 1) —but found no significant effect (pearson’s correlation test: r=0.004, p =0.98). one possible reason is that the priming effect is measured by comparison with the unrelated condition (rt on excellent given the movie is good vs. the movie is foreign.) the “unrelated” words (here, foreign) might introduce variation across items, which might obscure a potential by-item effect. future work should use a more uniform unrelated condition, e.g. identity (the movie is excellent) or antonym priming (the movie is bad). 8. conclusion. in this paper, we provide evidence from semantic priming that scalar alternatives (excellent) are activated in the processing of the relevant si-triggering sentences (the movie is good). this suggests that scalar alternatives pattern similarly to focus alternatives. at the same time, a number of empirical puzzles remain for future work, and we have argued for caution in using priming evidence to draw strong conclusions about pragmatic theory. proceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 238 https://doi.org/10.3765/elm https://www.elm-conference.net/ references bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67(1). 1–48. 10.18637/jss.v067.i01. bott, lewis & emmanuel chemla. 2016. shared and distinct mechanisms in deriving linguistic enrichment. journal of memory and language 91. 117–140. 10.1016/j.jml.2016.04.004. new approaches to structural priming. bott, lewis & steven frisson. 2022. salient alternatives facilitate implicatures. plos one 17(3). 1–10. 10.1371/journal.pone.0265781. braun, bettina & lara tagliapietra. 2009. the role of contrastive intonation contours in the retrieval of contextual alternatives. language and cognitive processes 25(7-9). 1024–1043. 10.1080/01690960903036836. de carvalho, alex, anne c. reboul, jean-baptiste van der henst, anne cheylus & tatjana nazir. 2016. scalar implicatures: the psychological reality of scales. frontiers in psychology 7. 10.3389/fpsyg.2016.01500. fraundorf, scott h., aaron s. benjamin & duane g. watson. 2013. what happened (and what did not): discourse constraints on encoding of plausible alternatives. journal of memory and language 69(3). 196–227. 10.1016/j.jml.2013.06.003. fraundorf, scott h., duane g. watson & aaron s. benjamin. 2010. recognition memory reveals just how contrastive contrastive accenting really is. journal of memory and language 63(3). 367–386. 10.1016/j.jml.2010.06.004. gotzner, nicole & jacopo romoli. 2022. meaning and alternatives. annual review of linguistics 8(1). 213–234. 10.1146/annurev-linguistics-031220-012013. gotzner, nicole & katharina spalek. 2017. role of contrastive and noncontrastive associates in the interpretation of focus particles. discourse processes 54(8). 638–654. 10.1080/0163853x.2016.1148981. gotzner, nicole, isabell wartenburger & katharina spalek. 2016. the impact of focus particles on the recognition and rejection of contrastive alternatives. language and cognition 8(1). 59–95. 10.1017/langcog.2015.25. grice, herbert paul. 1967. logic and conversation. in paul grice (ed.), studies in the way of words, 41–58. harvard university press. horn, laurence r. 1972. on the semantic properties of logical operators in english: ucla dissertation. husband, e. matthew & fernanda ferreira. 2015. the role of selection in the comprehension of focus alternatives. language, cognition and neuroscience 31(2). 217–235. 10.1080/23273798.2015.1083113. katzir, roni. 2007. structurally-defined alternatives. linguistics and philosophy 30(6). 669–690. 10.1007/s10988-008-9029-y. kim, christina s., christine gunlogson, michael k. tanenhaus & jeffrey t. runner. 2015. context-driven expectations about focus alternatives. cognition 139. 28–49. 10.1016/j.cognition.2015.02.009. lupker, stephen j. & penny m. pexman. 2010. making things difficult in lexical decision: the impact of pseudohomophones and transposed-letter nonwords on frequency and semantic primproceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 239 https://doi.org/10.3765/elm https://www.elm-conference.net/ ing effects. journal of experimental psychology: learning, memory, and cognition 36(5). 1267–1289. 10.1037/a0020125. rastle, kathleen, jonathan harrington & max coltheart. 2002. 358,534 nonwords: the arc nonword database. the quarterly journal of experimental psychology section a 55(4). 1339– 1362. 10.1080/02724980244000099. pmid: 12420998. rees, alice & lewis bott. 2018. the role of alternative salience in the derivation of scalar implicatures. cognition 176. 1–14. 10.1016/j.cognition.2018.02.024. repp, sophie & katharina spalek. 2021. the role of alternatives in language. frontiers in communication 6. 10.3389/fcomm.2021.682009. ronai, eszter & ming xiang. 2022. quantifying semantic and pragmatic effects on scalar diversity. in proceedings of the linguistic society of america 7(1), . ronai, eszter & ming xiang. to appear. three factors in explaining scalar diversity. in daniel gutzmann & sophie repp (eds.), proceedings of sinn und bedeutung 25, . rooth, mats. 1985. association with focus: university of massachusetts, amherst dissertation. rooth, mats. 1992. a theory of focus interpretation. natural language semantics 1(1). 75–116. 10.1007/bf02342617. sanford, alison j. s., jessica price & anthony j. sanford. 2009. enhancement and suppression effects resulting from information structuring in sentences. memory & cognition 37(6). 880– 888. 10.3758/mc.37.6.880. schwarz, florian, jérémy zehr, daniel grodner & hezekiah akiva bacovcin. 2016. subliminal priming of alternatives does not increase implicature responses. poster presented at the logic and language in conversation workshop, university of utrecht. spalek, katharina, nicole gotzner & isabell wartenburger. 2014. not only the apples: focus sensitive particles improve memory for information-structural alternatives. journal of memory and language 70. 68–84. 10.1016/j.jml.2013.09.001. sperber, dan & deirdre wilson. 1995. relevance: communication and cognition. wileyblackwell 2nd edn. swinney, david a. 1979. lexical access during sentence comprehension: (re)consideration of context effects. journal of verbal learning and verbal behavior 18(6). 645–659. 10.1016/s00225371(79)90355-4. swinney, david a, william onifer, penny prather & max hirshkowitz. 1979. semantic facilitation across sensory modalities in the processing of individual words and sentences. memory & cognition 7(3). 159–165. 10.3758/bf03197534. thomas, matthew a., james h. neely & patrick o’connor. 2012. when word identification gets tough, retrospective semantic processing comes to the rescue. journal of memory and language 66(4). 623–643. 10.1016/j.jml.2012.02.002. van tiel, bob, emiel van miltenburg, natalia zevakhina & bart geurts. 2016. scalar diversity. journal of semantics 33(1). 137–175. 10.1093/jos/ffu017. yan, mengzhu & sasha calhoun. 2019. priming effects of focus in mandarin chinese. frontiers in psychology 10. 10.3389/fpsyg.2019.01985. zehr, jeremy & florian schwarz. 2018. penncontroller for internet based experiments (ibex). https://doi.org/10.17605/osf.io/md832. proceedings of elm 2: 229-240, 2023 eszter ronai and ming xiang: tracking the activation of scalar alternatives with semantic priming. 240 https://doi.org/10.3765/elm https://www.elm-conference.net/ beyond surprising: english event structure in the maze lisa levinson* abstract. to what extent can we tease apart semantic representations and processes from other influences on processing such as probabilistic prediction? in this paper i detail two experiments testing the hypothesis that there is semantic complexity in the lexical representations of result verbs that influences reaction times above and beyond probabilistic distributions. this is done by replicating a self-paced reading study from levinson & brennan (2016) while also modelling lexical surprisal. experiment 1 replicates the original result, but only in experiment 2 using the maze task does the effect emerge beyond surprisals. the more focal maze task results suggest that processing costs associated with bieventive result verbs should be accounted for by grammatical factors, in addition to probabilistic prediction. 1. introduction. one of the central questions in experimental semantics is how to link the outcomes of various behavioral and neural tasks to the rules and representations of theoretical work. online tasks such as reaction times and eeg component amplitudes introduce a variety of other influences that must be teased apart from grammatical factors to identify a potentially causal association between a specific grammatical phenomenon and an observed effect. in this paper i explore such questions specifically within the domain of event structure, asking whether there are semantic, event structural properties in the lexical representations of verbs that influence reaction times above and beyond the probabilistic distribution of those verbs and their arguments. more specifically, in this study i test the hypothesis that event structure complexity is associated with a processing cost that cannot fully be explained by lexical surprisal, where surprisal is quantified using the transformer language model gpt-2 (radford et al. 2019). behavioral studies on a variety of event-structure related contrasts in verbs (discussed below) have found processing costs associated with greater complexity. many such studies have further argued that these costs bolster support for specific semantic event structures. however, many of these studies pre-date important advances in our understanding of the role of prediction in sentence processing and the development of language models which seem to more accurately model human-like predictions. in the spirit of delogu et al. (2017)’s work on complement coercion, in this paper i revisit the results of levinson & brennan (2016)’s experiment 2 on causative event structure in result verbs to test (a) whether the basic findings replicate, and (b) whether the effect of complex event structure goes beyond the effects expected due to sentence prediction, as modelled by surprisals from a transformer language model. this work is part of a series of replications in this vein to improve our understanding of the relationship between event structure, prediction, and sentence processing. previous behavioral studies have found “costs” for lexical semantic verb representations due to *much gratitude to ras yizhi tang, lila tappan, brighton pauli, yasemin gunal, emma thronson, and thea kendall-green who all made valuable contributions to this project! thanks also to the organizers and participants of elm 2. this work was supported by university of michigan’s urop program. author: lisa levinson, university of michigan (lisalev@umich.edu). proceedings of elm 2: 176-188, 2023 c©2023 lisa levinson published by the lsa with permission of the author(s) under a cc by license. 176 https://doi.org/10.3765/elm https://www.elm-conference.net/ the number of sub-events (mckoon & macfarland 2000, mckoon & macfarland 2002, gennari & poeppel 2003, mckoon & love 2011) and event types (gennari & poeppel 2003), even in lexical decision where contextual prediction does not play a role. it remains unclear, however, how these effects link with underlying semantic representations, and by what mechanisms they induce such costs. structural verb biases (such as frequency of transitive vs. intransitive frames) vary both within and across languages independent of the event structure of the verbs themselves (mckoon & macfarland 2000, rappaport hovav 2020). event structural properties thus might not travel through the same “causal bottleneck” (levy 2008) of surprisal, but rather make an independent contribution to processing. the majority of prior findings cannot tease apart these factors; while based on stimuli that are controlled for a variety of probabilistic factors, they have not been recently re-evaluated in the context of (a) probabilities calculated with less sparse language models, (b) measures such as surprisal that are more closely correlated with reading times (hale 2001, 2016), (c) statistical modeling of multiple stimulus co-variates, and (d) more focal behavioral tasks such as grammatical maze (freedman & forster 1985). 2. background. levinson & brennan (2016) tested the hypothesis that even phonologically identical verb forms could show varying processing costs depending on the event complexity associated with their syntactic context. for example, with “result” verbs such as melt (called “causative” verbs in the original paper), the same verb form may denote a bieventive causative change-of-state (1), or a monoeventive change-of-state (2). (1) the sun melted the ice. (2) the ice melted. transitives in this alternation in english have been analyzed in various studies as denoting more complex events than intransitives (inchoatives) (alexiadou et al. 2006, pylkkänen 2008, rappaport hovav & levin 2012). although the details of the number and types of events vary across specific implementations, the hypothesis being tested is whether there is a processing cost due to a greater number of subevents (for example, 2 subevents) in the transitive result verb as compared with it’s intransitive variants, which have been more commonly analyzed as denoting a simple event. in order to control for the potential confound of number of arguments, the study included sentences that also vary in number of arguments, but not number of subevents. these are “manner” verbs (described as “activity” verbs in the original study): (3) the student read the book. (4) the student read. the predicted (and observed) effect of event structure was thus an interaction such that transitive result verbs as in (1) are associated with longer reading times than intransitive result verbs (2) to a greater degree than transitive manner verbs (3) might take longer than intransitive manner proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 177 https://doi.org/10.3765/elm https://www.elm-conference.net/ verbs (4). 2.1. estimating human prediction. although it had been well established for many years that the probability of a word given prior context plays an important role in word and sentence processing, hale (2001) proposed the idea of quantifying this in terms of cognitive load measured as word surprisal. for example, given the partial sentence or prefix “after work, i like to”, one can extract from a predictive language model the conditional probability of that string, and then the probability of that string plus another word, such as “sleep”. as described in more detail in hale (2016), these prefix probabilities can then be mathematically transformed into surprisal values. surprisal inverts the scale so that higher values are less expected, and also log transforms them. these transformed values derived from conditional probabilities have been shown to correlate well with various behavioral and neurolinguistic measures, including self-paced reading (levy 2008). surprisal values can be computed from conditional probabilities derived from a variety of sources, from cloze task (taylor 1953) measurements to corpus-based bigram and large transformer models. smith & levy (2011) demonstrate that cloze probabilities differ significantly from corpus statistics. michaelov et al. (2021) further show that even when measuring all predictions in surprisal, computational language models more accurately predict the human n400 response, an eeg component strongly associated with expectation and prediction during sentence comprehension. put together, this work suggests that the best way to currently model the predictions that humans are making during sentence processing is with probabilities extracted from language models and quantified as surprisals. in the paper being replicated here, several standard measures were taken to attempt to control for non-event-structure influences on sentence processing, both in stimuli creation and the statistical analysis. these included frequency estimates derived from a corpus and transitivity biases estimated by ngram searches of google books, and norming to match sentence acceptability across conditions. however these lexicaland sentence-level statistics do not provide very direct insight into the probability or predictability of the verb and following words in the specific sentence context. for example, the verb “knit” is assigned a 66% transitivity probability by gahl et al. (2004). however, when presented with the sentence context “after work, i like to knit”, only 3 of the top 10 gpt-2 predictions are transitive continuations. 2.2. tasks for investigating the role of prediction. prior behavioral work on event structure has used tasks such as lexical decision (gennari & poeppel 2003), whole sentence reading times (mckoon & macfarland 2000, mckoon & macfarland 2002), self-paced reading (gennari & poeppel 2003, brennan & pylkkänen 2010), and a variant of self-paced reading called the stopmaking sense task (mckoon & love 2011). these methods provide various challenges to teasing apart the role of prediction from other influences on the outcome variables. single word lexical decision and whole-sentence reading don’t provide the context and granularity to explore questions of prediction. self-paced reading is more granular, but spillover effects make it difficult to separate responses to a current word from those for the prior word. the stop making sense task (mauner et al. 1995) requires participants to make a deeper decision when proceeding to the next word, as they are asked to specifically detect anomalies when the sentence stops making sense. this task seems to have potential to address the challenge of spillover and incrementality, depending on the proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 178 https://doi.org/10.3765/elm https://www.elm-conference.net/ design of “sense” violations (semantic, syntactic, or lexical) and the number and difficulty of distractor sentences. however, few studies have been conducted using this paradigm thus far. another task which addresses these challenges in a similar fashion but has a larger literature demonstrating its outcomes and comparison with other tasks is the maze task (freedman & forster 1985). the maze task is detailed in forster et al. (2009) as an alternative to self-paced reading that leads the participant to more incremental processing. in this task, rather than pressing a button to advance to the next word in the sentence, the participant must first choose the best sentence continuation from two presented choices. the outcome variable is still the response time for the next word in the sentence, but this response now requires an additional discrimination between possibilities. an example of subsequent screens in maze task presentation is given in figure 1. figure 1: maze task frame sequence the distractor alternatives can be designed to be “incorrect” continuations on various dimensions. in the grammatical maze, or g-maze, the alternatives are intended to be ungrammatical continuations. for example, in the context of the english words “that zebra chases”, the word “approves” would be an ungrammatical continuation, and might serve as a distractor for the grammatical option of the same length, “leopards”. in the lexical maze, or l-maze, the distractors are not known words, so the distractor might be a pseudoword such as “bunklet”. although the maze task is less naturalistic than self-paced reading, it has several apparent benefits for studying certain phenomena that justify this trade-off. forster et al. (2009) showed that the task is sensitive to syntactic complexity, and also detail advantages such as lack of spillover effects, and the necessity for participants to commit to a parse at each decision point. this in turn ensures that words are being more fully integrated prior to each button press, rather than delayed to a later point in the sentence. they also found robust effects similar in several respects to eye tracking, though with more incrementality due to the inability to return to earlier portions of the text. boyce et al. (2020) further demonstrated that the task can be successfully implemented in web-based studies, and that it has greater statistical power and effect “localization” (to a word in the sentence) than even in-lab self-paced reading. 3. experiment 1. experiment 1 sought to directly replicate effects of crossing event complexity and transitivity in english (experiment 2 of levinson & brennan 2016), using the same method proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 179 https://doi.org/10.3765/elm https://www.elm-conference.net/ table 1: stimuli for experiments 1 and 2 (from @levinson.brennan2016a) exx w1 w2 det n v v+1 v+2 w8 args verb type 1 what did the cook thaw in the cafeteria? 2 result 2 when did the popsicle thaw in the cafeteria? 1 result 3 what did the teacher hum for the students? 2 manner 4 when did the teacher hum for the students? 1 manner of self-paced reading, with additional predictors in the statistical analyses to evaluate the relative contribution of event complexity vs. surprisal. the replication was also conducted online rather than in-lab (the setting of the original study). 3.1. method. 3.1.1. participants. 90 american english readers completed the study. participants were undergraduate students at oakland university and were compensated with course credit. 3.1.2. materials. the stimuli were the same as those used in experiment 2 of levinson & brennan (2016), where conditions were normed and matched for acceptability but subject animacy varies to allow for this matching. verb frequency was also matched across verbtype conditions. as seen in table 1, the conditions cross verb type (result vs. manner) with number of arguments.1 more details about the stimuli design can be found in the original paper. there were 87 experimental sentence pairs (43 result, 44 manner). the fillers used were different from those used in the original study, but followed a similar pattern. 36 were ungrammatical questions, to balance out the experimental items for the acceptability task. there were 60 additional declarative sentence fillers with varying ratios of ungrammatical sentences per participant, depending on the stimuli list. for the experimental materials, a latin square design was used to create two lists such that no participants would see both items in a verb pair. 3.1.3. gpt-2 surprisals. surprisal for each non-initial word in the sentences was estimated using probabilities generated by the open source transformer language model gpt-2 (radford et al. 2019), as shown for the regions surrounding the verb in figure 2. these surprisals were calculated in python using the minicons package (misra 2022), which provides convenience wrappers for the hugging face transformers library (wolf et al. 2020). 1full stimuli are available via the osf repository: https://osf.io/zh3ub/. proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 180 https://doi.org/10.3765/elm https://www.elm-conference.net/ n v v+1 manner result manner result manner result 0 5 10 15 verb type g p t − 2 s ur pr is al ( bi ts ) arguments 1 2 gpt−2 surprisals figure 2: surprisals for each item by word position from the pre-verbal noun to first word following the verb. crossbars show mean and standard error of the mean. 3.1.4. procedure. participants completed a self-paced moving window task (with postsentence acceptability judgments) presented online using ibex presentation software hosted on ibexfarm (drummond 2013). participants used the space bar to reveal the sentence word-by-word at their own pace. after the final word in the sentence was presented, a question mark appeared. the subject was instructed to respond press the ‘f’ key if the sentence was natural, and the ‘j’ key if the sentence was unnatural. before the main experiment, participants completed 10 practice trials with feedback for acceptability judgment accuracy. 3.1.5. data analysis. prior to analysis, trials that were judged to be unacceptable were removed. participants and items with mean accuracy below 70% on acceptability judgments were also excluded six participants (out of 90, 7%) and 12 experimental items (out of 87, 14%). no trials were excluded on the basis of reading times, as analyzing log transformed reading times minimized the impact of outliers. the critical regions analyzed were the verb and the spillover region at verb+1 (preposition). a linear mixed effects models (gelman & hill 2006, baayen et al. 2008) was fit at the verb position, with a model similar to that used in the replicated study, which i will call the “event structure” model (5). this model included fixed effects for the interaction between verbtype and causativity, using treatment coding for the categorical predictors (with ‘manner’ as the reference level for verbtype, and ‘intransitive’ for transitivity). also included were fixed effects for centered and scaled lexical statistics and random intercepts for participants and items.2 frequency was sourced 2models with random slopes did not converge consistently across different models. bayesian models with equivalent formulas plus random slopes for verbtype, transitivity, and surprisal (where relevant) were fit with flat priors in brms (bürkner 2017) with stan (team 2022) to serve as a pseudo-maximum likelihood estimate that is most comparable to those fit by lme4. the coefficients from these models did not substantively differ from those fit by lme4. proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 181 https://doi.org/10.3765/elm https://www.elm-conference.net/ from the corpus of contemporary american english (coca) (davies 2008). for the v+1 spillover region, where significant effects were observed in the original study, a full model was also fit including a fixed effect of gpt-2 surprisals, as in (6). finally, a “surprisal” model was fit omitting the event structure interaction (7). (5) log(rt) ~ verbtype * transitivity + frequency + length + (1 | participants) + (1| items) (6) log(rt) ~ verbtype * transitivity + verb_surprisal + verb_frequency + verb_length + (1 | participants) + (1|items) (7) log(rt) ~ verb_surprisal + verb_frequency + verb_length + (1 | participants) + (1|items) since the effect in the spillover region would be associated with the verb, and not the matched post-verbal words, the surprisals, frequency, and length factored into the analysis of spillover regions were for the verb, not v+1 (a very short preposition). models were fit using the lme4 package (bates & maechler 2009) in r (r development core team 2006), and p-values estimated using the implementation of satterthwaite approximation provided by the lmertest package (kuznetsova et al. 2017). 3.2. results. the results are visualized word-by-word for the verb and two words following in figure 3. as in the original study, there was no significant interaction at the verb. results for the spillover region supported replication of the predicted interaction in the event structure model (β = .03, se = .018, p = .046). a posthoc pairwise comparison showed a significant effect for transitivity in the result verbs as well (β = .03, se = .01, p = .009). however, when verb surprisal was added to the model (full model), only surprisal evidenced a significant effect (β = .02, se = .008, p = .028). while model comparison between the event structure model and full model showed that gpt-2 significantly improved model fit (likelihood ratio test, p = .025), the addition of event structure in the full model did not improve fit over the surprisal model. proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 182 https://doi.org/10.3765/elm https://www.elm-conference.net/ n v v+1 manner result manner result manner result 300 1000 3000 verb type r ea ct io n t im e (m s, lo g ax is s ca le ) arguments 1 2 self−paced reading results (experiment 1) figure 3: mean participant reaction times by word position from the pre-verbal noun to first word following the verb. crossbars show mean and standard error of the mean. 3.3. experiment 1 discussion. the critical interaction effect at the verb+1 region from levinson & brennan (2016) did replicate, in that a similar model omitting the predictor of surprisal did provide support for an effect of event structure in transitive result verbs that was not present in the manner verbs. however, once surprisal was added to the full model, it seemed to absorb the effect attributed to event structure. this does not mean that event structure does not play a role, or that the analysis of result or manner verbs is incorrect. crucially, surprisal is not a measure that is completely independent from semantic representations. the predictions generated by language models are built from probability distributions that are themselves a product of human language production. thus, as discussed above, surprisal is a “bottleneck” through which semantic and other grammatical representations are transformed into predictions. that said, the results of experiment 1 are also consistent with the hypothesis that grammatical contrasts in the stimuli do, to some extent, evade the surprisal bottleneck and more directly influence sentence processing. in self-paced reading, this influence may be obscured by the dispersion of processing costs across multiple words in the spillover region. 4. experiment 2. experiment 2 was designed using the maze task in order to encourage participants to read the stimuli sentences more incrementally and carefully, in order to minimize the dispersion of the influence of specific lexical items and maximize the effect at the verb. concentrating the verb-triggered processing to the verb response time should allow for a more focalized comparison between the role of semantic and other grammatical word properties and probabilistic measures such as surprisal. 4.1. method. proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 183 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4.1.1. participants. 60 participants with english as their first language and living in the united states were recruited via prolific. payment was $3.33 for approximately 20 minutes to complete the study, for an hourly rate of $10. ages ranged from 18 to 81, with a mean age of 35. 4.1.2. materials. the experimental stimuli sentences were the same as those in experiment 1, but word alternatives were also generated for the maze task. these alternatives were generated using the a-maze package (boyce et al. 2020) with the included pre-trained grnn language model (gulordava et al. 2018). after automatic generation, the alternatives were manually checked and adjusted. where the alternative did not seem sufficiently ungrammatical or inappropriate, it was changed to another alternative with the same word length that was deemed “incorrect” based on experimenter intuition (as all maze alternatives were generated in the original studies using this method). since the same materials were used for both experiments, stimuli measures such as gpt-2 surprisals and frequencies are also identical to those gathered for experiment 1. since there is no offline acceptability judgment in the maze task, all fillers were grammatical sentences. the experimental stimuli were split into 4 lists (rather than 2) to reduce study completion time, and each participant saw 43 or 44 experimental stimuli plus 40 grammatical declarative sentence fillers. 4.1.3. procedure. the experiment was run online via ibexfarm using the ibex stimuli generation scripts from the a-maze package (boyce et al. 2020).3 to implement the maze task, each sentence is presented in a series of frames. each frame shows two sentence continuation candidates for the participate to choose from, to the left and right of center. the correct continuation is randomly assigned to a side and varies for individual items across participants to counterbalance. since there is no “continuation” at the first word, the first alternative is simply “x-x-x”. if a participant made the correct selection, the next frame would be displayed. if they chose the alternative, they were presented with error feedback and the rest of that item was not presented. the time from the presentation of the frame until their selection is recorded and serves as the primary outcome variable, the response time. the session started with instructions followed by 3 practice trials with positive and negative feedback. for the experimental materials, a latin square design was used to create lists where no participants would see both items in a verb pair. participants were given a 30-second break every 12 sentences. 4.1.4. data analysis. all trials with incorrect maze responses were excluded from the analysis. since trials ended whenever an incorrect choice was made, subsequent words in the same sentence were not presented. the same range of event structure, surprisal, and full models were fit as for the self-paced reading results, but only at the verb since there was no evidence of spillover effects across the maze experiment. for these data, it was also possible to add a random slope of surprisals for participants (in models including surprisal) and transitivity for participants, so the full model was as in (8). 3although the scripts use the original ibex controllers and the maze controller developed by boyce et al. (2020), a clonable demo of the experiment is hosted on the pcibexfarm (schwarz & zehr 2021) and linked from the osf repository: https://osf.io/zh3ub/. proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 184 https://doi.org/10.3765/elm https://www.elm-conference.net/ (8) log(rt) ~ verbtype * transitivity + verb_surprisal + verb_frequency + verb_length + (1 + verb_surprisal | participants) + (1 + transitivity |items) 4.2. results. as predicted, the maze results exhibited more focal effects, with no apparent spillover, as can be seen in figure 4. even with the full model (including surprisal along with event structure predictors), the interaction (β = .13, se = .04, p = .002) was significant at the verb. the predicted pairwise effect for transitivity in causative verbs was not significant in the full model (β = .08, se = .05, p = .07), but a likelihood ratio test comparing that model to one omitting transitivity suggested significantly improved fit from the event structure predictors (p < .001). model comparison between the full model and partial models showed that both event structure predictors (lrt p < .001) and surprisals (lrt p < .001) improved model fit over either event structure or surprisals “alone”. noun verb verb+1 manner result manner result manner result 1000 3000 5000 verb type r t s (m s, lo g ax is s ca le ) arguments 1 2 grammatical maze results (experiment 2) figure 4: mean participant reaction times by word position from the pre-verbal noun to first word following the verb. crossbars show mean and standard error of the mean. 4.3. experiment 2 discussion. effects of event complexity beyond surprisal are evident in the maze task data. this not only supports the hypothesis that some event structural properties evade the surprisal bottleneck, but also demonstrates that a more incremental task such as maze can help to tease apart these variables. 5. conclusion. in conclusion, these results support an independent contribution of event structure complexity to incremental processing above and beyond surprisal in the slower but more incremental maze task. comparison of methods suggests that such effects may only be separable with more focal and larger effects that allow for teasing apart multiple fine-grained contributions to sentence processing. references. alexiadou, artemis, elena anagnostopoulou & florian schäfer. 2006. the properties of antiproceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 185 https://doi.org/10.3765/elm https://www.elm-conference.net/ causatives crosslinguistically. in henk van riemsdijk, harry van der hulst, jan koster & mara frascarelli (eds.), phases of interpretation, vol. 91, 187–212. berlin, new york: mouton de gruyter. doi:10.1515/9783110197723.4.187. baayen, r. harald, douglas j. davidson & douglas m. bates. 2008. mixed-effects modeling with crossed random effects for subjects and items. journal of memory and language 59(4). 390–412. doi:10.1016/j.jml.2007.12.005. bates, douglas & martin maechler. 2009. lme4: linear mixed-effects models using s4 classes. boyce, veronica, richard futrell & roger p. levy. 2020. maze made easy: better and easier measurement of incremental processing difficulty. journal of memory and language 111. 104082. doi:10.1016/j.jml.2019.104082. brennan, jonathan & liina pylkkänen. 2010. processing psych verbs: behavioral and meg measures of two different types of semantic complexity. language and cognitive processes 25(6). 777–807. doi:10.1080/01690961003616840. bürkner, paul-christian. 2017. brms: an r package for bayesian multilevel models using stan. journal of statistical software 80. 1–28. doi:10.18637/jss.v080.i01. davies, mark. 2008. word frequency data from the corpus of contemporary american english (coca). https://www.wordfrequency.info. delogu, francesca, matthew w. crocker & heiner drenhaus. 2017. teasing apart coercion and surprisal: evidence from eye-movements and erps. cognition 161. 46–59. doi:10.1016/j. cognition.2016.12.017. drummond, alex. 2013. ibex farm. forster, kenneth i., christine guerrera & lisa elliot. 2009. the maze task: measuring forced incremental sentence processing time. behavior research methods 41(1). 163–171. doi:10. 3758/brm.41.1.163. freedman, sandra e. & kenneth i. forster. 1985. the psychological status of overgenerated sentences. cognition 19(2). 101–131. doi:10.1016/0010-0277(85)90015-0. gahl, susanne, dan jurafsky & douglas roland. 2004. verb subcategorization frequencies: american english corpus data, methodological studies, and cross-corpus comparisons. behavior research methods, instruments, & computers 36(3). 432–443. doi:10.3758/bf03195591. gelman, andrew & jennifer hill. 2006. data analysis using regression and multilevel/hierarchical models. cambridge university press. gennari, silvia & david poeppel. 2003. processing correlates of lexical semantic complexity. cognition 89(1). b27–b41. doi:10.1016/s0010-0277(03)00069-6. gulordava, kristina, piotr bojanowski, edouard grave, tal linzen & marco baroni. 2018. colorless green recurrent networks dream hierarchically. arxiv:1803.11138 [cs] . hale, john. 2001. a probabilistic earley parser as a psycholinguistic model. in second meeting of the north american chapter of the association for computational linguistics, doi:10.3115/ 1073336.1073357. hale, john. 2016. information-theoretical complexity metrics. language and linguistics compass 10(9). 397–412. doi:10.1111/lnc3.12196. kuznetsova, alexandra, per b. brockhoff & rune h. b. christensen. 2017. lmertest package: tests in linear mixed effects models. journal of statistical software 82. 1–26. doi:10.18637/jss. proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 186 https://doi.org/10.3765/elm https://www.elm-conference.net/ v082.i13. levinson, lisa & jonathan brennan. 2016. the costs of zero-derived causativity in english: evidence from reading times and meg. in daniel siddiqi & heidi harley (eds.), morphological metatheory, vol. 229 linguistik aktuell/linguistics today, 163–198. john benjamins. levy, roger. 2008. expectation-based syntactic comprehension. cognition 106(3). 1126–1177. doi:10.1016/j.cognition.2007.05.006. mauner, gail, michael k. tanenhaus & gregory n. carlson. 1995. implicit arguments in sentence processing. journal of memory and language 34(3). 357–382. doi:10.1006/jmla.1995.1016. mckoon, gail & jessica love. 2011. verbs in the lexicon: why is hitting easier than breaking? language and cognition 3. 313–330. doi:10.1515/langcog.2011.011. mckoon, gail & talke macfarland. 2000. externally and internally caused change of state verbs. language 76(4). 833–858. doi:10.2307/417201. mckoon, gail & talke macfarland. 2002. event templates in the lexical representations of verbs. cognitive psychology 44. doi:10.1016/s0010-0285(02)00004-x. michaelov, james a., seana coulson & benjamin k. bergen. 2021. so cloze yet so far: n400 amplitude is better predicted by distributional information than human predictability judgements. arxiv:2109.01226 [cs, math] . misra, kanishka. 2022. minicons: enabling flexible behavioral and representational analyses of transformer language models. arxiv:2203.13112 [cs] . pylkkänen, liina. 2008. introducing arguments. cambridge, ma: mit press. r development core team. 2006. r: a language and environment for statistical computing. vienna, austria: r foundation for statistical computing. radford, alec, jeffrey wu, rewon child, david luan, dario amodei, ilya sutskever et al. 2019. language models are unsupervised multitask learners. openai blog 1(8). 9. rappaport hovav, malka. 2020. deconstructing internal causation. in elitzur a. bar-asher siegal & nora boneh (eds.), perspectives on causation: selected papers from the jerusalem 2017 workshop jerusalem studies in philosophy and history of science, 219–255. cham: springer international publishing. rappaport hovav, malka & beth levin. 2012. lexicon uniformity and the causative alternation. in the theta system, oxford: oxford university press. doi:10.1093/acprof:oso/9780199602513. 003.0006. schwarz, florian & jeremy zehr. 2021. tutorial: introduction to pcibex – an open-science platform for online experiments: design, data-collection and code-sharing. proceedings of the annual meeting of the cognitive science society 43(43). smith, nathaniel & roger levy. 2011. cloze but no cigar: the complex relationship between cloze, corpus, and subjective probabilities in language processing. in proceedings of the annual meeting of the cognitive science society, vol. 33, . taylor, wilson l. 1953. “cloze procedure”: a new tool for measuring readability. journalism quarterly 30(4). 415–433. doi:10.1177/107769905303000401. team, stan development. 2022. stan modeling language users guide and reference manual, version 2.3. wolf, thomas, lysandre debut, victor sanh, julien chaumond, clement delangue, anthony moi, proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 187 https://doi.org/10.3765/elm https://www.elm-conference.net/ pierric cistac, tim rault, rémi louf, morgan funtowicz, joe davison, sam shleifer, patrick von platen, clara ma, yacine jernite, julien plu, canwen xu, teven le scao, sylvain gugger, mariama drame, quentin lhoest & alexander m. rush. 2020. transformers: state-of-the-art natural language processing. in proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, 38–45. online: association for computational linguistics. proceedings of elm 2: 176-188, 2023 lisa levinson: beyond surprising: english event structure in the maze. 188 https://doi.org/10.3765/elm https://www.elm-conference.net/ representing affect information in word embeddings yuhan zhang, wenqi chen, ruihan zhang & xiajie zhang* abstract. a growing body of research in natural language processing (nlp) and natural language understanding (nlu) is investigating human-like knowledge learned or encoded in the word embeddings from large language models. this is a step towards understanding what knowledge language models capture that resembles human understanding of language and communication. here, we investigated whether and how the affect meaning of a word (i.e., valence, arousal, dominance) is encoded in word embeddings pre-trained in large neural networks. we used the human-labeled dataset (mohammad 2018) as the ground truth and performed various correlational and classification tests on four types of word embeddings. the embeddings varied in being static or contextualized, and how much affect specific information was prioritized during the pre-training and fine-tuning phase. our analyses show that word embedding from the vanilla bert model (devlin et al. 2019) did not saliently encode the affect information of english words. only when the bert model was fine-tuned on emotion related tasks or contained extra contextualized information from emotion-rich contexts could the corresponding embedding encode more relevant affect information. keywords. language models; word embeddings; affect meaning; lexical semantics 1. introduction. with the success of large neural network models in completing complicated language tasks, evaluating the models’ interpretability and intrinsic capabilities has become a heated research trend (e.g., manning et al. 2020, mikolov et al. 2013). the evaluation work could be roughly classified into two types: one that relies on the output of the language models (lms) to infer the model’s linguistic ability and the other that looks into the components of lms (e.g., word embeddings) for such inspiration. while previous evaluation tasks have focused on testing lms’ explicit linguistic knowledge (e.g., syntactic knowledge such as islands, semantic knowledge such as compositionality, word-level knowledge such as polysemy), we pick a piece of knowledge that is less studied but essential to intelligence. specifically, we studied whether and how word embeddings learned via supervised methods in large neural networks encode the affect information of a word (e.g., valence, arousal, dominance). our work shows that even though contextualized word embeddings were in general better at capturing intricate affect meanings compared to static word embeddings, especially after being fine-tuned on emotion related tasks, word embeddings from vanilla bert did not attain salient affect knowledge. in section 2, we detailed relevant work that led us to our investigation. in section 3, we laid out the unsupervised and supervised methodologies we took and in section 4, we spelled out the findings. in section 5, we discussed implications of our research and possible future directions. *we would like to thank mycal tucker, roger levy, and the audience at elm 2022 for their generous feedback. all mistakes are ours. authors: yuhan zhang, harvard university (yuz551@g.harvard.edu); wenqi chen, harvard university (wenqichen@g.harvard.edu); ruihan zhang, massachusetts institute of technology (ruihanz@mit.edu); xiajie zhang, massachusetts institute of technology (xiajie@mit.edu). proceedings of elm 2: 310-321, 2023 c©2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang published by the lsa with permission of the author(s) under a cc by license. 310 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2. related work. three major aspects in the current nlp world are related to our work: the linguistic knowledge revealed to be grasped by language models, the relationship between what word embeddings have achieved to represent and the actual lexical semantics, and how affect information is studied in natural language processing. 2.1. what language models can learn. a growing body of research has investigated what linguistic knowledge large artificial neural networks can learn while they are trained in a supervised way to predict the next word. in the syntactic category, scholars have shown that large lms have varying grammatical knowledge ranging from the local subject-verb agreement to long-distance filler-gap dependencies (e.g., hu et al. 2020, linzen et al. 2016, warstadt et al. 2020, wilcox et al. 2018). aside from looking at predicted words to infer lms’ linguistic knowledge, manning et al. (2020) show that a linear transformation of word embeddings from bert (devlin et al. 2019) captures linguistic hierarchical structures. there is also increasing attention to understanding lms’ abilities to represent meanings. in semantics and pragmatics, promising and positive results seem to support lms’ increasingly sophisticated abilities such as doing natural language inference (e.g., poliak et al. 2018, wang et al. 2018). instead of relying on direct output of lms, finding the relationship between human understanding of language and representations in word embeddings have also been fruitful. 2.2. word embedding and lexical semantics. since deep contextual language models provide contextualized word embeddings which naturally encode the distance between word tokens in a vector space, this property can be utilized to test whether the trained distance in word embeddings reflects the natural way words group together according to our lexical semantic knowledge. existing studies have shown that the pre-trained bert model is able to place polysemous words that appear in different contexts into distinct regions of the shared vector space (wiedemann et al. 2019) and the word sense distances are correlated with human judgments (nair et al. 2020). there is also the general claim that contextualized word embedding in bert is good at word sense disambiguation (loureiro et al. 2021). but in addition to this line of research on word sense disambiguation, investigation into other aspects of lexical semantics is limited. we made an attempt in this work to bring together studies of word embedding representations and another aspect of lexically encoded meaning – the affect information – as a novel case study of what lms can learn from supervised training. 2.3. affect information in nlp. affect information refers to any explicit or implicit emotion related information. we use the word affect instead of emotion because we do not rely on common emotion words such as happy and sad as baselines of comparison; rather, we rely on three primary independent dimensions of emotions as scales to quantify the meaning of emotion. the three dimensions are valence (positiveness-negativeness/pleasantness-unpleasantness), arousal (activepassive, some people also take it to mean the intensity of the emotion invoked by the word), and dominance (dominant-submissive, or the level of control exerted by the word) (osgood et al. 1957, russell 1980, 2003). another reason to choose this terminology is due to its validity and consistency: there are already multiple human-labeled datasets that are based on these scales as the quantitative ground truth (bradley & lang 1999, mohammad 2018, warriner et al. 2013). in the fields that are related to emotion recognition, sentiment analysis, and affective comproceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 311 https://doi.org/10.3765/elm https://www.elm-conference.net/ puting, transformer-based models and more fine-grained annotation of emotion categories have gathered the attention of the nlp/nlu community (e.g., alswaidan & menai 2020, demszky et al. 2020, rashkin et al. 2018, suresh & ong 2021). zooming into existing studies relevant to word embeddings and affect information, it has been shown that multi-dimensional embeddings of emotional words achieved higher performance in sentiment analysis tasks than human annotated emotion vectors (li et al. 2017). furthermore, the multi-dimensional word embeddings for emotional words are often added onto base static word embeddings for downstream sentiment analysis tasks (mao et al. 2019). while existing research focused more on the application of emotional word embeddings onto sentiment analysis tasks, seldom has explored whether state-of-the-art contextualized word embeddings have learned the affect meaning of all kinds of words. our study is set up to fill the gap. 3. methodology. our methods focus on constructing the relationship between word embeddings and human cognition. on the human cognition side, we adapt the dataset from the affect ratings of humans from cognitive studies. on the word embedding side, we adopted both contextualized and static word embeddings. in order to probe the relationships, we take both the unsupervised and supervised methods to reflect the relationships. 3.1. affect ratings of human as the ground truth. we selected the nrc vad dataset (national research councial (canada) valence arousal dominance) curated by mohammad (2018) as the ground truth for affect meanings of english words.1 there are 20,000 english words annotated along the three independent dimensions of affect. in each dimension, the rating starts from 0 to 1 where 0 means the most negative, passive, or submissive in the v, a, or d dimension and 1 means the most positive, active, or dominant in their respective dimension. this is the largest manually created vad corpus in any language. 3.2. word embeddings under study. in this paper, we studied four kinds of word embeddings as described below. we took from each word embedding the same subset of words that are compatible with the nrc vad corpus. the embeddings differ from each other by how much and what kinds of contextualized information are captured from training. we took glove as a representative of existing static word embeddings in the nlp community and compared it with more contextualized ones. we also varied the kinds of contextualized information by prioritizing emotion related and sentiment analysis related information. • pre-trained word embeddings using the glove algorithm (pennington et al. 2014) which we refer to as the glove embedding in this paper. the word embeddings were downloaded from stanford nlp website2 which were trained on wikipedia and gigaword 5 data. • pre-trained word embeddings that were retrieved from the last hidden layer of the base bert model (devlin et al. 2019) after we directly ran the model over the nrc word list. we refer to this type of word embedding representation as base bert throughout the paper. • pre-trained word embeddings retrieved from the last hidden layer of the bert model fined tuned on the goemotion dataset (demszky et al. 2020). the goemotion dataset contains 58k english reddit comments labeled by humans on 27 emotion categories. we refer to this 1data can be downloaded from http://saifmohammad.com/webpages/nrc-vad.html 2https://nlp.stanford.edu/projects/glove/ proceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 312 https://doi.org/10.3765/elm https://www.elm-conference.net/ as goemotion bert in the paper. • given that the bert-based word embeddings are supposed to be contextualized because they could be functions of the entire input sentence (ethayarajh 2019), we utilized this contextual sensitivity feature and output contextualized bert word embeddings by deriving them from running imdb movie reviews (maas et al. 2011). for each nrc word, we first extracted all its embeddings by retrieving the last hidden layer of the bert model, we then computed the first principal component to derive the single word vector for each individual word. 3.3. unsupervised probe: principal component analysis. firstly, in order to investigate whether the types of word embeddings under investigation encode affect information, we conducted dimensionality reduction via principal component analysis (pca from the sklearn package (buitinck et al. 2013, pedregosa et al. 2011)) over each of the studied word embeddings. pca is a convenient tool for such visualization without us going deep into deciphering the architecture of multi-dimensional word embeddings. specifically, we sought for the correlation between each principal component from the embeddings and the human ratings of each vad dimension. higher correlation indicates that the high-dimensional information in that word embedding is more likely to saliently encode affect information. this unsupervised approach provides a rudimentary insight into whether affect meanings are well captured by the word embeddings being investigated. 3.4. unsupervised probe: cosine distance for semantic similarity. it has been shown that cosine distances between vector representations of words indicate the words’ semantic relatedness (dumais et al. 1988, mikolov et al. 2013, bojanowski et al. 2017). the literature we learned from is nair et al. (2020). they investigated whether contextualized word embeddings – bert embeddings in this case – capture human-like distinctions between english word senses, such as polysemy (chicken as animal vs. chicken as meat) and homonymy (bat as mammal vs. bat as sports equipment). for each pair of target word senses, they measured (1) the cosine distance of the two embeddings and (2) the human judgment of the relatedness of the word senses in a 2-dimensional spatial arrangement task. they then applied spearman rank correlation analysis to these two measurements. the results turn out that the distance of word senses in bert’s embedding space correlated with human judgments and that the correlation of homonyms was higher. learning from nair et al. (2020), we took a small sample of affect words and ran correlation tests to see the relationship between the word embedding space and the vector space of human judgments (i.e., the vad 3-dimensional space). similar to the pca results, a stronger correlation means a more salient encoding of real-world affect information in the investigated word embeddings. more about the sampling method: shaver et al. (1987) defined six basic emotion categories (i.e., love, surprise, sadness, anger, joy, fear) and a list of affect words in each category, totalling 132 affect words. we took 80 affect words from shaver et al. (1987) (e.g., disgust, envy, enjoyment, desire) and calculated the cosine similarity between each and another word, resulting in 80 × 79 similarity scores. we did the same thing by iterating over the kind of vector space, from the nrc vad 3-dimensional human judgment space, to the four kinds of word embeddings introduced in section 3.2, resulting in 80 × 79 × 5 pairwise similarity scores. then, we did spearman rank correlation analysis between one and another embedding representations. we adopted a nonparametric correlation like nair et al. (2020) because, first, the numerical distribution of a given type proceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 313 https://doi.org/10.3765/elm https://www.elm-conference.net/ of word embedding does not satisfy the normality assumption and, second, we cared more about whether the two embedding representations share a monotonic relationship, instead of a strictly linear relationship. 3.5. supervised probe: linear classifier. in addition, we experimented with supervised learning methods to probe whether the word embeddings can be linearly separable into the vad dimensions. we trained a one-layer neural network (i.e., a logistic regression model) to predict the binary vad labels3 with the embeddings as inputs. the linear classifier was designed to be as simple as possible to reveal information encoded in the embeddings. if the prediction results match well with the human criterion, it is reasonable to conclude that the word embeddings have the capability to represent the affective information. 4. analyses & results. 4.1. principal component analysis. figure 1 shows the pca result. we extracted 5,586 words that are in the vocabulary of glove, base bert, goemotion bert, and contextualized bert embeddings, ran pca over these word embeddings, and represented the numerical value of the first two principal components along the two axes. each horizontal panel represents one dimension of vad. each column in figure 1 represents a type of tested word embedding. each point is color coded with human judgment. the darker the color, the higher the rating is along one of the vad dimensions. table 1 shows the spearman correlation coefficient between the numerical values of human judgments along each affect dimension and the numerical values of the corresponding principal component given a type of word embedding. from table 1, we learn that every type of word embedding can capture some aspects of the 3-dimensional affect meaning based on the significance level of the correlation coefficients. yet judging from the magnitude of correlation coefficients for both principal components as well as the number of correlation coefficients that are larger than 0.2 out of the three vad dimensions, we see that the goemotion bert embedding and the contextualized bert embedding outperform the other two. dimension pca glove base bert goemotion bert contextualized bert valence 1st 0.117 (.000) 0.015 (.247) 0.084 (.000) -0.077 (.000) 2nd -0.006 (.672) 0.133 (.000) 0.227 (.000) 0.291 (.000) arousal 1st -0.005 (.696) 0.011 (.412) 0.285 (.000) 0.251 (.000) 2nd 0.238 (.000) -0.094 (.000) -0.042 (.002) -0.007 (.599) dominance 1st 0.316 (.000) 0.028 (.037) 0.160 (.000) 0.086 (.000) 2nd 0.182 (.000) 0.111 (.000) 0.283 (.000) 0.380 (.000) table 1: spearman correlation coefficients (p values in the bracket) between human ratings (01) in each affect dimension and each corresponding principal component of four types of word embeddings (correlation coefficients in bold when larger than 0.2, also see each subplot in figure 1) in table 2, we present the explained variance ratios to show what percentage of the variance in the whole embedding space was explained by the first two principal components for each type 3we transformed the numerical human rating (0-1) to a binary categorical variable with the threshold of 0.5. proceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 314 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: 2-components pca representations of glove, base bert, goemotion bert, and contextualized bert & human ratings for valence, arousal, and dominance (human ratings were color-coded in a spectrum. each dot represents a word. 5,586 words are represented. the darker the dot, the more prominent the human rating in the respective dimension.) of word embedding. two pieces of observations are worth interpreting. first, the relatively low ratio of explained variance across the eight values indicates that the information represented in the first two principal components might not be representative to generalize on the capacity of certain word embedding to encode certain affect meaning. there might be some affect meaning that has been captured in higher dimensions of the vector space but not projected in the pca measure. we would need more sensitive probes other than pca in future investigations. second, assuming that the pca analysis is a great approximant for affect encoding, the fact that the first component of base bert accounts for 29.4% of the total variance and yet does not correlate with vad as much as the other word embeddings shows that base bert might not be optimized to distinguish intricate affect meaning of words. this offers a nice window to peek into what linguistic information is and isn’t weighted the most for bert. explained ratio glove base bert goemotion bert contextualized bert first 0.040 0.294 0.118 0.021 second 0.030 0.040 0.070 0.016 table 2: explained variance ratio of the two principal components proceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 315 https://doi.org/10.3765/elm https://www.elm-conference.net/ 4.2. cosine similarity. table 3 captures the spearman correlation coefficients based on pairwise cosine similarity of 80 words whose embeddings were from each type of word embeddings. the main comparison is in the first column where human judgments on the semantic similarity of the 80× 79 pairs of words are compared with the cosine distance of those word pairs based on different types of word embedding. the result shows that all the four types of word embeddings show a significant correlation with human affect judgment. the base bert embedding has the weakest correlation of all, which also echoes with figure 1. then, in the order of more relatedness is the goemotion bert embedding, glove, and the contextualized bert. it is worthwhile to deliberate on why the static word embeddings like glove outperforms the base bert. also, it is interesting that even though the goemotion bert is fined tuned on emotion specific task, the correlation is still worse than bert-based embeddings with movie specific contextualized information. corr (p) vad glove base goemotion contextualized vad 1.000 (.000) glove 0.272 (.000) 1.000 (.000) base 0.116 (.000) 0.148 (.000) 1.000 (.000) goemotion 0.252 (.000) 0.172 (.000) 0.013 (.471) 1.000 (.000) contextualized 0.314 (.000) 0.710 (.000) 0.240 (.000) 0.204 (.000) 1.000 (.000) table 3: spearman correlation (p value) of pairwise cosine similarities between each and the rest of word embedding types (coefficients larger than 0.2 and less than 1 are in bold.) 4.3. linear classifier. in this section, we report the classification results from the linear classifier probe. the training covariate matrix of the classification model was each type of word embeddings. the response variables were the binary categorical label representing lower or higher range of one of the vad dimensions. the conversion was based on a 0.5 threshold. each type of word embedding with each affect dimension constituted a linear classifier and together we constructed 12 classifiers. we used the nrc vad vocabulary as the training and the validation data (with a 70/30 split). the test data comprised of 130 word with strong affect information randomly extracted from shaver et al. (1987). we compared the performance of glove, base bert, goemotion bert, and contextualized bert embeddings on each of the three affect dimensions and present the result in table 4. it is clear that the contexualized bert embedding shows superior performance compared to the others. to be noted, our validation set included many words whose vad score might not be as salient and intuitive as the ones in the test set. this might explain that the validation set received lower accuracy score than the test sample. besides, we observed that all of the model embeddings are best at predicting the valence dimension compared with arousal and dominance. this matches our prior knowledge because valence represents the positive or negative of the words and carries the most straightforward meaning for the context to capture. but arousal and dominance will be more complicated and subtle to capture depending on its distributional semantics. interestingly, goemotion bert does not perform very well in the dimensions of arousal and dominance. this phenomenon leaves ample space to investigate further how word embeddings are trained to capture certain meaning but not the others. proceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 316 https://doi.org/10.3765/elm https://www.elm-conference.net/ embedding types validation accuracy affect word sample accuracy valence arousal dominance valence arousal dominance glove 0.75 0.70 0.73 0.84 0.74 0.75 base bert 0.76 0.73 0.74 0.93 0.85 0.82 goemotion bert 0.68 0.68 0.70 0.92 0.76 0.76 contextualized bert 0.85 0.77 0.85 0.95 0.88 0.90 table 4: performance of different word embeddings in predicting vad labels. 5. discussion & conclusion. in this work, we investigated the capability of word embeddings to capture affect meaning. by combining results from unsupervised and supervised learning, we found that our contextualized bert embedding, where the representation of each word is the first principal component of all embeddings of the same word in its multiple occurrences in the imdb dataset. in comparison, the base bert does not show prominent encoding of affect meaning and sometimes is even inferior to the static glove. the goemotion bert embedding also does not perform as well as the contextualized bert embedding. out of the vad dimensions, it is hard to know which dimension is relatively well represented or easy to capture: based on the pca results, all vad dimensions could attain a correlation higher than 0.2 in some embedding representations; based on the linear classifier results, valence seems to be better represented than the other two. the discrepant pattern might result from the different sample sizes used in these tasks (i.e., 5,586 words for pca and 130 words as the test set for the linear classifiers) and more controlled experiments are needed to pin down the fact. 5.1. implications. here we provide a novel aspect of testing word embedding capabilities that involve intricate emotion and affect related meaning. based on what lexical and linguistic knowledge word embeddings can and cannot capture, especially for the renowned bert embedding (ettinger 2020, klafka & ettinger 2020, manning et al. 2020, pandia et al. 2021), we add one additional piece of information. the relatively poor encoding of affect information in bert might suggest these lines of implications: (1) the affect information of english lexicon is by nature implicit and hard to capture from the distributional features of words in huge corpora; (2) for sentiment related nlp tasks or affect computing tasks (maas et al. 2011, socher et al. 2013, zhang et al. 2015), we might consider using affect enriched embeddings or leveraging additional emotional capabilities of lms from transfer learning. 5.2. limitations & future work. as far as we know, this is the first attempt in the nlp community to investigate the capability of word embeddings to encode affect information in words. we lay three lines of research for future investigations. first, from an algorithmic point of view, we would like to ask what specific training mechanisms enable word embeddings to capture certain information, especially those intricate meanings that don’t rely on co-occurrence and distributional semantics. second, talking about meaning and semantics in a broader sense, there are still tons of unknowns that fall outside of the realm of formal semantics or compositionality. more rigorous identifications and descriptions of different kinds of meaning are needed for studying the knowledge of meaning in lms. third, to what extent can language carry affective information? this is an essential question in affective computing (picard 2000, 2003, tao & tan 2005). a better underproceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 317 https://doi.org/10.3765/elm https://www.elm-conference.net/ standing from both the algorithmic level and the linguistic level of affect encoding and recognition would provide valuable insights into this question. references alswaidan, nourah & mohamed el bachir menai. 2020. a survey of state-of-the-art approaches for emotion recognition in text. knowledge and information systems 62(8). 2937–2987. bojanowski, piotr, edouard grave, armand joulin & tomas mikolov. 2017. enriching word vectors with subword information. acl transactions of the association for computational linguistics 5. 135–146. bradley, margaret m & peter j lang. 1999. affective norms for english words (anew): instruction manual and affective ratings. tech. rep. the center for research in psychophysiology. buitinck, lars, gilles louppe, mathieu blondel, fabian pedregosa, andreas mueller, olivier grisel, vlad niculae, peter prettenhofer, alexandre gramfort, jaques grobler, robert layton, jake vanderplas, arnaud joly, brian holt & gaël varoquaux. 2013. api design for machine learning software: experiences from the scikit-learn project. in ecml pkdd workshop: languages for data mining and machine learning, 108–122. demszky, dorottya, dana movshovitz-attias, jeongwoo ko, alan cowen, gaurav nemade & sujith ravi. 2020. goemotions: a dataset of fine-grained emotions. in proceedings of the 58th annual meeting of the association for computational linguistics (acl), 4040– 4054. online: association for computational linguistics. 10.18653/v1/2020.acl-main.372. https://aclanthology.org/2020.acl-main.372. devlin, jacob, ming-wei chang, kenton lee & kristina toutanova. 2019. bert: pre-training of deep bidirectional transformers for language understanding. in proceedings of the 2019 conference of the north american chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), 4171–4186. minneapolis, minnesota: association for computational linguistics. 10.18653/v1/n19-1423. https: //aclanthology.org/n19-1423. dumais, susan t, george w furnas, thomas k landauer, scott deerwester & richard harshman. 1988. using latent semantic analysis to improve access to textual information. in proceedings of the sigchi conference on human factors in computing systems, 281–285. ethayarajh, kawin. 2019. how contextual are contextualized word representations? comparing the geometry of bert, elmo, and gpt-2 embeddings. in proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlp-ijcnlp), 55–65. hong kong, china: association for computational linguistics. 10.18653/v1/d19-1006. https: //aclanthology.org/d19-1006. ettinger, allyson. 2020. what bert is not: lessons from a new suite of psycholinguistic diagnostics for language models. transactions of the association for computational linguistics (acl) 8. 34–48. hu, jennifer, jon gauthier, peng qian, ethan wilcox & roger levy. 2020. a systematic assessment of syntactic generalization in neural language models. in proceedings of the 58th annual meeting of the association for computational linguistics (acl), 1725–1744. onproceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 318 https://doi.org/10.3765/elm https://www.elm-conference.net/ line: association for computational linguistics. 10.18653/v1/2020.acl-main.158. https: //aclanthology.org/2020.acl-main.158. klafka, josef & allyson ettinger. 2020. spying on your neighbors: fine-grained probing of contextual embeddings for information about surrounding words. arxiv preprint arxiv:2005.01810 . li, minglei, qin lu, yunfei long & lin gui. 2017. inferring affective meanings of words from word embedding. ieee transactions on affective computing 8(4). 443–456. 10.1109/taffc.2017.2723012. linzen, tal, emmanuel dupoux & yoav goldberg. 2016. assessing the ability of lstms to learn syntax-sensitive dependencies. transactions of the association for computational linguistics 4. 521–535. loureiro, daniel, kiamehr rezaee, mohammad taher pilehvar & jose camacho-collados. 2021. analysis and evaluation of language models for word sense disambiguation. computational linguistics 47(2). 387–443. maas, andrew l., raymond e. daly, peter t. pham, dan huang, andrew y. ng & christopher potts. 2011. learning word vectors for sentiment analysis. in proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies (acl), 142–150. portland, oregon, usa: association for computational linguistics. http: //www.aclweb.org/anthology/p11-1015. manning, christopher d, kevin clark, john hewitt, urvashi khandelwal & omer levy. 2020. emergent linguistic structure in artificial neural networks trained by selfsupervision. proceedings of the national academy of sciences 117(48). 30046–30054. https://doi.org/10.1073/pnas.1907367117. mao, xingliang, shuai chang, jinjing shi, fangfang li & ronghua shi. 2019. sentiment-aware word embedding for emotion classification. applied sciences 9(7). 10.3390/app9071334. https://www.mdpi.com/2076-3417/9/7/1334. mikolov, tomas, ilya sutskever, kai chen, greg s corrado & jeff dean. 2013. distributed representations of words and phrases and their compositionality. in advances in neural information processing systems (neurips), 3111–3119. mohammad, saif. 2018. obtaining reliable human ratings of valence, arousal, and dominance for 20,000 english words. in proceedings of the 56th annual meeting of the association for computational linguistics (volume 1: long papers), 174–184. melbourne, australia: association for computational linguistics. 10.18653/v1/p18-1017. https://aclanthology.org/ p18-1017. nair, sathvik, mahesh srinivasan & stephan c. meylan. 2020. contextualized word embeddings encode aspects of human-like word sense knowledge. proceedings of the cognitive aspects of the lexicon workshop at the 28th international conference on computational linguistics (coling) https://arxiv.org/abs/2010.13057. osgood, charles egerton, george j suci & percy h tannenbaum. 1957. the measurement of meaning 47. university of illinois press. pandia, lalchand, yan cong & allyson ettinger. 2021. pragmatic competence of pre-trained language models through the lens of discourse connectives. arxiv preprint:2109.12951 . proceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 319 https://doi.org/10.3765/elm https://www.elm-conference.net/ pedregosa, f., g. varoquaux, a. gramfort, v. michel, b. thirion, o. grisel, m. blondel, p. prettenhofer, r. weiss, v. dubourg, j. vanderplas, a. passos, d. cournapeau, m. brucher, m. perrot & e. duchesnay. 2011. scikit-learn: machine learning in python. journal of machine learning research 12. 2825–2830. pennington, jeffrey, richard socher & christopher d. manning. 2014. glove: global vectors for word representation. in empirical methods in natural language processing (emnlp), 1532–1543. http://www.aclweb.org/anthology/d14-1162. picard, rosalind w. 2000. affective computing. mit press. picard, rosalind w. 2003. affective computing: challenges. international journal of humancomputer studies 59(1-2). 55–64. poliak, adam, aparajita haldar, rachel rudinger, j. edward hu, ellie pavlick, aaron steven white & benjamin van durme. 2018. collecting diverse natural language inference problems for sentence representation evaluation. in proceedings of the 2018 emnlp workshop blackboxnlp: analyzing and interpreting neural networks for nlp, 337–340. brussels, belgium: association for computational linguistics. 10.18653/v1/w18-5441. https: //aclanthology.org/w18-5441. rashkin, hannah, eric michael smith, margaret li & y-lan boureau. 2018. towards empathetic open-domain conversation models: a new benchmark and dataset. arxiv preprint arxiv:1811.00207 . russell, james a. 1980. a circumplex model of affect. journal of personality and social psychology 39(6). 1161. russell, james a. 2003. core affect and the psychological construction of emotion. psychological review 110(1). 145. shaver, phillip, judith schwartz, donald kirson & cary o’connor. 1987. emotion knowledge: further exploration of a prototype approach. journal of personality and social psychology 52(6). 1061. socher, richard, alex perelygin, jean wu, jason chuang, christopher d manning, andrew y ng & christopher potts. 2013. recursive deep models for semantic compositionality over a sentiment treebank. in proceedings of the 2013 conference on empirical methods in natural language processing, 1631–1642. suresh, varsha & desmond c ong. 2021. using knowledge-embedded attention to augment pretrained language models for fine-grained emotion recognition. in 2021 9th international conference on affective computing and intelligent interaction (acii), 1–8. ieee. tao, jianhua & tieniu tan. 2005. affective computing: a review. in international conference on affective computing and intelligent interaction, 981–995. springer. wang, alex, amanpreet singh, julian michael, felix hill, omer levy & samuel bowman. 2018. glue: a multi-task benchmark and analysis platform for natural language understanding. in proceedings of the 2018 emnlp workshop blackboxnlp: analyzing and interpreting neural networks for nlp, 353–355. brussels, belgium: association for computational linguistics. 10.18653/v1/w18-5446. https://aclanthology.org/w18-5446. warriner, amy beth, victor kuperman & marc brysbaert. 2013. norms of valence, arousal, and dominance for 13,915 english lemmas. behavior research methods 45(4). 1191–1207. proceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 320 https://doi.org/10.3765/elm https://www.elm-conference.net/ warstadt, alex, alicia parrish, haokun liu, anhad mohananey, wei peng, sheng-fu wang & samuel r bowman. 2020. blimp: the benchmark of linguistic minimal pairs for english. transactions of the association for computational linguistics 8. 377–392. wiedemann, gregor, steffen remus, avi chawla & chris biemann. 2019. does bert make any sense? interpretable word sense disambiguation with contextualized embeddings. conference on natural language processing http://arxiv.org/abs/1909.10430. wilcox, ethan, roger levy, takashi morita & richard futrell. 2018. what do rnn language models learn about filler–gap dependencies? in proceedings of the 2018 emnlp workshop blackboxnlp: analyzing and interpreting neural networks for nlp, 211–221. brussels, belgium: association for computational linguistics. 10.18653/v1/w18-5423. https://aclanthology.org/w18-5423. zhang, xiang, junbo zhao & yann lecun. 2015. character-level convolutional networks for text classification. advances in neural information processing systems (neurips ) 28. proceedings of elm 2: 310-321, 2023 yuhan zhang, wenqi chen, ruihan zhang and xiajie zhang: representing affect information in word embeddings. 321 https://doi.org/10.3765/elm https://www.elm-conference.net/ on the interpretation of german einige. the effect of tense and cardinality maya cortez espinoza & lea fricke* abstract. we present a study investigating the effect of tense (past vs. future) on the computation of scalar implicatures in connection with the german quantifier einige ‘some’ in an interactive experiment, which included a financial incentive for participants to consider whether another speaker would share their judgment. we tested the hypothesis that scalar implicatures are less frequently drawn in future tense than in past tense. in addition, we studied to what extent sets with various cardinalities are prototypical representatives of einige + n. we hypothesized that larger cardinalities are more prototypical representatives of the quantifier einige than smaller cardinalities (relative to the cardinality of the total set). we analyzed the experimental data with probabilistic bayesian models with a linking hypothesis between participants’ responses and readings based on utility maximization in simple decision problems. in line with the hypotheses, we found that less scalar implicatures are drawn in future tense than in past tense, which replicates the results of previous research on english some, and that with an increase in set size acceptance of statements involving einige also increases. keywords. scalar implicatures; german einige; tense; quantifier interpretation; bayesian modelling; experimental pragmatics; 1. introduction. a sentence like (1-a), involving the quantifier some, usually triggers the scalar implicature (si) (1-b). (1) a. some students danced. b. si: not all students danced. a factor that influences the computation of sis, but that has not yet received much attention in the literature, is tense, to the effect that sis are less frequently drawn in future tense than in past tense. normally, the gricean reasoning process upon hearing a statement involving the quantifier some, like (1-a), goes roughly as follows: under the assumption that the speaker is cooperative and follows the maxim of quantity, a hearer assumes that she chose some as the maximally informative expression and that she cannot truthfully use the stronger expression all. thus, the hearer concludes that not all students danced.1 moreover, as observed by geurts (2010), the assumption that the speaker is maximally informed about the issue at hand (“competence assumption”) is crucial for drawing the implicature. not being able to truthfully assert the stronger alternative all students danced is compatible with *we would like to thank edgar onea, daniele panizza, swantje tönnis, and alexandre cremers for their insights in fruitful disscusions and helpful suggestions. we further thank simon dampfhofer, kathrin hirsch, kassandra keinz, lea-sophie kravanja, michaela leeb, melanie loitzl, and eva winkler for their help conducting this experiment. this research was funded by linglab graz, which we thankfully acknowlegde. authors: maya cortez espinoza, karlfranzens-universität graz (maya.cortez-espinoza@uni-graz.at) & lea fricke, karl-franzens-universität graz/ruhruniversität bochum (lea.fricke@rub.de). 1alternatively, sis can be derived from applying a silent logical operator on various sites (fox 2007, chierchia 2013). with this account, sis can be derived both at the utterance level and in embedded positions. proceedings of elm 2: 49-60, 2023 c©2023 maya cortez espinoza and lea fricke published by the lsa with permission of the author(s) under a cc by license. 49 https://doi.org/10.3765/elm https://www.elm-conference.net/ not being opinionated about the stronger alternative. thus, to take the “epistemic step” (sauerland 2004) and conclude that the speaker implicated that not all students danced, the hearer needs to assume that the speaker is competent about the issue. as the future is inherently uncertain, a statement about the future, like (2) is a mere prediction. for this reason, it makes sense for the speaker to use a weak expression to make her statement semantically compatible with both outcomes (some or all students danced). on considering this, the hearer will not take the epistemic step and no implicature will be drawn. (2) some students will dance. this effect of tense has to our knowledge only once been investigated for english and italian in an acquisition experiment by chierchia et al. (1998)2. the authors argue that a prediction about the future constitutes a context in which an si is suspended. in the experiment, children had to judge the truth of statements made by a puppet, which contained one of the scalar expressions some and or. two contexts were tested. in the first context, which chierchia et al. (1998) label ‘description mode’, the target sentence is in past tense and in the second, which is called ‘prediction mode’, it is in future tense. with regard to some, they report preliminary results according to which there was 14% and 34% acceptance for the si violation in english and italian respectively in the description mode and more than 75% acceptance in both languages in the prediction mode. the final results for or are similar: in the description mode, the acceptance rate for the si-violation is lower than 33% in both languages while it is at ceiling in the prediction mode. the results of this study support the hypothesis that tense affects the computation of sis and that sis are less frequently drawn in future tense than in past tense. however, its limitations are that there are no data for adult speakers and that in the case of some only preliminary data are reported. we aim to replicate these results for german einige ‘some’ with adult speakers. for the interpretation of the quantifier einige or some, another factor beside the sis comes into play, namely the question which cardinalities can be felicitously described with this quantifier. according to the set-theoretic definition, a statement involving some, like (1-a), is true iff the intersection between the set of dancers and the set of students is non-empty (e.g. barwise & cooper 1981). however, previous research suggests that the exact cardinality of the intersection matters to people, i.e. some cardinalities are more typical representatives of some than others. van tiel & geurts (2014) studied the interpretation of quantifiers in english. they presented participants with pictures of ten circles, of which varying numbers were black, together with the sentence quantifier of the circles are black. participants had to either judge the truth value of the statement or rate on a likert scale how well the statement describes the situation (typicality). they found that for the quantifier ‘some’, the numbers 2–4 received the highest truth ratings while the numbers 4–6 received the highest typicality rating. another study investigating the processing of some, degen & tanenhaus (2015), included naturalness ratings of a statement like “you got some of the gumballs” with contexts of varying numbers from a total of 13 gumballs. additionally to judging naturalness, participants could judge a statement as false. the authors found that statements in which some referred to the set sizes 1, 2, 3 or 13 gumballs were rated as less natural than statements in which some referred to sets of 6–8 2we thank an anonymous reviewer for pointing out this paper to us. proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 50 https://doi.org/10.3765/elm https://www.elm-conference.net/ gumballs. when some referred to a set of one gumball, the statement was rated as false in 12% of the cases. they also found that particularly for small but also for large set sizes, exact number terms were rated as more natural than some. we note that in both studies numbers that roughly represent half of the total sample receive the highest naturalness/typicality ratings and that it appears to be odd to refer to a set size of 1 with some. responsible for the latter effect seems to be the plural marking of the noun phrase which leads to a plurality inference, as discussed in spector (2007). in our study, we compare the acceptability of statements involving einige referring to different set sizes ranging from 1–9. note that german has two different versions of ‘some’, manche and einige. according to our intuitions, manche is associated with smaller numbers compared to einige. additionally, when stressed einige has a prominent reading of ‘many’, as example (3) shows. (3) a: wie viele leute haben bereits gegessen? (‘how many people have eaten already’?) b: naja, schon einige. (‘well, quite a few’.) thus, we want to test the hypothesis that statements involving einige will receive higher acceptance when referring to sets with higher cardinalities, relative to the cardinality of the total set, than when referring to sets with smaller cardinalites. a further goal of our experiment was to control for the kind of reasoning induced by the experimental task. experiments with a classic truth-value judgement task design often only test for the general availability of a reading but do not control for the status of that interpretation with the consequence that it is not clear whether participants’ judgments are based on a semantic interpretation or whether other factors, such as communicative relevance, were considered. previous research has shown that the nature of the task crucially affects the rate of drawing implicatures. for example, with regard to embedded implicatures, geurts & pouscoulous (2009) showed that an inference task lead to a higher implicature rate compared to a truth-value-judgement task. benz & gotzner (2014) suggested that relevance of the si in the context of the experiment is a decisive factor for its computation. in our experiment, we employed the method developed by fricke et al. (2022). it aims to target interpretations that are of communicative relevance, that is, interpretations that the participant deems likely to be shared by another language user. to induce this kind of recursive thinking in participants, the design includes a monetary incentive: participants were told that they would lose money if another person did not share their judgment. this way, every single response had direct financial consequences for the participant, and it was in their own interest to think carefully about each item. the linking hypothesis is based on utility theory for simple decision problems: participants aim to maximize their expected utility measured in terms of financial payoff.3 2. experiment. we investigate the following hypotheses in our experiment: • the computation of sis of german einige is affected differently by future tense and past tense to the effect that the rate of implicatures is higher in past tense. 3we are aware that utility theory has received criticism, e.g. by tversky (1975) and kahneman & tversky (1979). however, in our experiment the translation between decisions and financial benefits is quite simple. moreover, we introduced some level of control regarding the effect of bias in participants’ decisions and found no such effect. proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 51 https://doi.org/10.3765/elm https://www.elm-conference.net/ • larger cardinalities are more prototypical representatives of the quantifier einige than smaller numbers and thus statements involving the quantifier einige will receive higher acceptance rates when referring to larger sets than to smaller sets. • the set of cardinality 1 is a particularly bad representative of the quantifier einige due to the plurality inference. 2.1. participants. we tested 32 native speakers of german (mostly austrian german) (mean age = 23.8 years, sd = 5.5 years, 15 female and 17 male), all of which were university students or former university students. they were recruited via a university-newsletter and received a financial compensation varying between 8.50c and 11.20c, with a mean of 10.18 and a standard deviation of 0.56, depending on their performance on control items. 2.2. materials. the target sentences were conditional statements containing the scalar term einige. they were presented in the context of a story about nine candidates in a reality show, who did activities together. the stimuli had the form of bets, made by a person named lina about activities to happen on the show and the participants’ task was to decide whether the bets were won (called ‘accepting the bet’ in the following) or lost (‘rejecting the bet’), thereby judging the truth of the target sentences. we manipulated three factors. the first factor was cardinality, which represents the number of candidates that were involved in an activity and which ranged from 0 to 9. the 0-context yields a false target sentence, the 9-context yields an si-violation and the numbers 1–8 constitute different manifestations of the quantifier einige. cardinality was tested within participants and within items. the second factor was tense. the verb of the target sentences was either in past tense (perfekt)4 or in future tense (futur i). this factor was tested at the participant level. participants were assigned either past tense or future tense. (half of the participants saw bets in past tense and the other half in future tense.) for past tense, people were told that the show had been prerecorded already before lina placed the bets but aired later, and therefore, lina worded her bets in past tense. for future tense, lina placed her bets before the recording of the show and therefore, bets were worded in future tense. the third factor, role, was manipulated at the participant level as well (meaning half of the participants acted in role 1 and the other half in role 2). in role 1, participants had to decide whether or not to redeem bets at a betting office. they received 5c starter cash. redeeming a bet cost a fee of 10 cents each. for each accepted bet that was actually won the participants received 30 cents payout. thus, in effect, the participants gained 20 cents for an accepted bet which was won, and they lost 10 cents for an accepted bet that was lost. in this role, a participant profited financially from accepted bets. to avoid a bias towards accepting borderline cases in this role, we installed the redeeming fee, with which participants would lose money when randomly submitting bets and would therefore consider whether a betting office agent might share their judgment. in role 2, participants acted as a betting office agent. they had to decide for each redeemed bet whether it was won or not. participants received 15c starter cash. for an accepted bet, they had to pay out 20 cents, while their budget remained unchanged when they rejected a bet. also, participants were told that lina, the person who had placed the bets, would raise an objection if a bet that was in fact 4in colloquial german, perfect tense is the default tense to talk about the past, see e.g. klis et al. (2017). proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 52 https://doi.org/10.3765/elm https://www.elm-conference.net/ won had been rejected. this would lead to a monetary loss of 30 cents. thus, participants had to consider whether their judgement would be shared by lina. in sum, in role 2, participants profited from rejected bets. to avoid a bias towards rejecting bets, we installed the financial deduction for incorrect decisions. these conditions for the two roles are summarized in table 1. role status of bet accepted bet rejected bet role 1 in fact won +20ct 0ct in fact lost -10ct 0ct role 2 in fact won -20ct -30ct in fact lost -20ct 0ct table 1: payoff table note that in calculating the compensation that a participant received, decisions on test items were always considered correct; only mistakes on fillers affected the final compensation negatively. however, participants did not know this before the experiment. to sum up, the factorial design was 10 cardinality x 2 tense x 2 role. there were two lexicalizations per cardinality, which varied between lists. test items were distributed over 16 experimental lists (half of them in future, half of them in past tense) with 20 test items and 2 participants per list. in addition to the test items, each list contained 32 filler items. 12 of of them were test items from a different experiment, and 9 uncontroversially won and 9 uncontroversially lost bets served as controls. figure 1 is an example of a test item for future tense. the stimuli were shown along with a table which indicated for each candidate whether she was involved in the activity in question. the numbers of involved candidates ranged from 0 to 9 (cardinality). additional context resolved the antecedent of the conditional as true. note that target sentences had the form of conditional sentences because we had an additional hypothesis about upward and downward entailing environments that we will not report on due to a procedure error, which happened on the level of participant instructions and which turned part of the data invalid. 2.3. procedure. the experiment took place in the lab of the theoretical and empirical linguistics research group at the university of graz. before the start of the actual experiment, participants saw 5 training items. they a) helped to get the participants accustomed to the task and b) made clear how to handle the conditional construction, which included the target sentence. then, depending on the role they were assigned, participants each received 5 or 15c starter cash in stacks of 10 and 20 cent coins. the participant saw the betting slips one at a time in a randomized order and had to decide whether to accept (pay out/redeem) the bet or not. if they wanted to accept, they had to return the betting slip together with the money to the experimenter. the experimenter entered the participant’s decision into an excel sheet that automatically calculated the sum the participant received as financial compensation after the experiment. the participant did not receive any feedback as to their gains and losses from individual bets, neither during nor after the experiment. after the first half of the betting slips, there was a short break. the experiment took 25 to 40 minutes in total. proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 53 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: sample item, cardinality 7, future tense 2.4. results. 2.4.1. data exclusion. the data of one participant (role 1, past tense) was excluded due to erring three out of 18 times on control items. therefore, data from 31 participants was subjected to further analysis (7 for the combination of role 1 and past tense, and 8 for all other tense x role combinations). 2.4.2. descriptive statistics. figure 2 shows the acceptance rates of bets by cardinality and tense. in the cardinality-0 context, in which no candidate was involved in the activity, acceptance is at 0 in both tenses. acceptance in the cardinality-9 context, which constitutes a si violation, differs between the tenses. in the past tense, acceptance is at 60% while it is at 93.8% in future tense. turning to the acceptance rates for cardinalities that represent einige ‘some’, we observe a) acceptance for cardinality 1 is particularly low (0/13.3%) and b) that acceptance increases with increasing cardinality. moreover, acceptance rates for cardinalities 1–4 differ between the tenses; they are higher in past tense than in future tense. our descriptive data shows that the two roles hardly differ in acceptance rates, as can be seen in table 2. therefore, we did not consider the factor role in the bayesian modelling. role won: count lost: count 1 108 (72.0%) 42 (28.0%) 2 115 (71.9%) 45 (28.1%) table 2: accepting rates between roles proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 54 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2: percentage of accepted bet by tense and cardinality 3. bayesian modelling. 3.1. concept. we created a bayesian model based on the assumption that a participant aims to maximize her utility for each decision in the experiment. in this model, a participant’s expected utility to act (uact), shown in (4) is the sum of the products of i) the subjective probability they assign the decision to be right (prs act, true) and the payoff for being right and ii) the subjective probability they assign to the decision being wrong (prs act, false) and the payoff for being wrong. prs act, true is different for the act of accepting and rejecting, since prs reject, true = 1 prs accept, true. (4) uact = payoffact, true ∗ prs act,true + payoffact, false ∗ (1− prs act, true) the probability for accepting a bet is then determined from the overall utility (the difference between uaccept and ureject) as log odds: (5) p(accept) = logistic((uaccept − ureject)) positive overall utilities, therefore, mean higher than 50% chance of accepting bets, negative utilities mean lower than 50% of accepting bets and utilities further away from 0 mean stronger tendencies to accept/reject. intuitively, therefore, strong prs and big rewards/punishments influence u greater than weak prs and small rewards/punishments. in our model, we do not assume differences in prs for different people as we did not have any hypotheses on some people acting differently than others. with the utility formula being constant across conditions, only subjective probabilities alternate. these conceptually depend on two factors: the computation of an si and the sensitivity of einige to different cardinalities. firstly, we assumed a general probability pr(si) for interpreting einige with the si not all. the alternative, pr(nosi), is the probability for interpreting einige without the si. as there are no other competing readings of einige, we assumed pr(nosi) = 1 − pr(si). our model assumes this probability to be equally accessible to all proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 55 https://doi.org/10.3765/elm https://www.elm-conference.net/ speakers of german. furthermore, our model assumes these probabilities to be different for the two tense levels: the probability to interpret einige without the si is higher in future tense than in past tense. of the two probabilities pr(si) and pr(nosi) either both, only one, or none of them contributes to prs accept, true, depending on whether the context supports the respective reading, see table 3. a participant’s prs accept, true is therefore the sum of her commitments to supported readings. this means that for cardinality 0, where none of the readings are supported, prs accept, true is set to 0, irrespective of the probabilities of the two readings. for cardinality 9, prs accept, true depends solely on pr(nosi), which enables us to tease apart the two readings’ probabilities. cardinality nosi-reading si-reading 0 not supported not supported 1–8 supported supported 9 supported not supported table 3: possible contexts and supported readings the second factor that contributed to prs act, true is cardinality sensitivity. for cardinalities 1– 8, both the nosiand the si-reading are supported and prs accept, true yields 1, but in these cases, cardinality sensitivity accounts for lower accepting rates. in our model, cardinality sensitivity is represented by a dirichlet distribution over cardinalities 1–8. the values of prototypicality are determined by the model itself through sampling. conceptually, we were interested in the prototypicality of all cardinalities relative to each other with the most prototypical cardinality having prs accept, true = 1. for this, each prototypicality value is divided by the highest prototypicality value. to sum up, table 4 shows all the possible cases for calculating prs accept, true. they depend on cardinality (c) and tense (t). case calculation of prs accept, true if c = 0: = 0 else if 1 ≤ c ≤ 8: = prototypicality(c) max(prototypicality) else if c = 9 and t = past: = pr(nosi)past else (c = 9 and t = future): = pr(nosi)future table 4: calculating prs accept, true 3.2. implementation and output. we conducted a mcmc simulation to construct our bayesian model in stan using the rstan package in r (stan development team 2022b,a, r core team 2022). we used uniform priors in order not to skew our results with previous hypotheses. when running the model, we used 15000 iterations, the first 5000 of which were dismissed as warm up. all models had excellent convergence as evidenced by visual checks and rhat values (1.00–1.01). after sampling, means were computed for each desired parameter. for model comparison, we used the bridgesampling package (gronau et al. 2020) and computed bayes factors proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 56 https://doi.org/10.3765/elm https://www.elm-conference.net/ (bf) for samples of 8000 iterations each. figure 3: probability density plot for pr(si) (with pr(nosi) = 1-pr(si)) as shown in table 5, the probability to interpret einige without the si is 13% higher for future tense than for past tense. pr(si) pr(nosi) past 0.38 0.62 future 0.25 0.75 table 5: si probabilities we additionally created a reduced model that does not distinguish between reading probabilities for tenses and compared it to our original model, which turned out superior to the reduced model (bf = 5.15). figure 3 shows a probability density plot for the movements of the parameter around the variable space during the simulation. the x-axis shows different probability values. for prototypicality values, see table 6 and the associated probability density plot in figure 4. cardinality 1 has the lowest prototypicality value with 0.05. cardinalities 2–7 are quite similar to each other with values ranging from 0.12 – 0.14, and cardinality 8 has the highest prototypicality value with 0.18. model comparison with a reduced model, in which no cardinality relative prototypicality was assumed, yielded that our original model is superior (bf > 150). cardinality 1 2 3 4 5 6 7 8 prototypicality 0.05 0.12 0.13 0.13 0.13 0.13 0.14 0.18 table 6: prototypicality values as our descriptive data suggested that there might be a difference for cardinality sensitivity between tenses (small numbers received less acceptance in future tense than in past tense), we proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 57 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 4: probability density plot for prototypicality values constructed another model that specified two different dirichlet distributions for prototypicality (future vs. past tense).5 model comparison showed that the original model was better than this more complex model (bayes factor > 150). 4. discussion. we found that sis are drawn less reliably in future tense than in past tense. the method we employed was intended to evoke rational behavior and targeted interpretations that are relevant from a communicative perspective. replicating the findings by chierchia et al. (1998) with this method suggests that the si-difference caused by tense that was found in that study is a robust effect and holds in communicative settings. as discussed in the introduction, this phenomenon can be explained pragmatically, namely with the epistemic step not being possible for future situations. however, in the case of our experiment, this explanation is not completely conclusive. in our design, the context was exactly the same for future tense and past tense. in both cases, the bet was made about an unknown situation. still, a clear difference was found between future tense and past tense. this hints at the effect of tense on sis being at least partly in the semantic domain – possibly the effect is a grammaticalized instance of the pragmatic principle described above. a further finding of our experiment is that the acceptability of statements involving einige increases with growing set sizes. this deviates from the findings in the literature on english some, according to which medium set sizes receive the highest ratings. this difference may stem from the fact that einige competes with another lexical item, manche, which is according to our intuitions associated with small numbers, at least in some cases.6 furthermore, stressed einige has a very 5note that we did not have a hypothesis about this. so this is a post hoc analysis. 6for manche to be felicitous, a basic set must be specified, which is not the case for einige. compare (i-a) in an out of the blue-context: (i) a. heute um sechs uhr haben sich einige/#manche leute am hasnerplatz verabredet. proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 58 https://doi.org/10.3765/elm https://www.elm-conference.net/ prominent reading of many, which is clearly not the case for english some. although, descriptively, we observed a tendency in the data for lower cardinalities to be less accepted in future tense than in past tense, this difference was not meaningful in the bayesian analysis. nevertheless, it would be interesting to gather more data, as an experiment targeted at this issue may find this to be significant. another idea for future research is testing whether there is a contrast between the rates at which sis are drawn between the english going-to-future and will-future. as the former is used to refer to future events that are more certain than the latter, we expect the future effect to be smaller for the going-to-future and the rate of drawn sis to be higher. references barwise, jon & robin cooper. 1981. generalized quantifiers and natural language. linguistics and philosophy 4(2). 159–219. https://doi.org/10.1007/bf00350139. benz, anton & nicole gotzner. 2014. embedded implicatures revisited: issues with the truth value judgement paradigm. in judith degen, michael franke & noah goodman (eds.), proceedings of the formal & experimental pragmatics workshop, 1–6. tübingen. chierchia, gennaro. 2013. logic in grammar: polarity, free choice, and intervention. oxford: oxford university press. https://doi.org/10.1093/acprof:oso/9780199697977.001.0001. chierchia, gennaro, stephen crain, maria teresa guasti & rosalind thornton. 1998. “some” and “or”: a study on the emergence of logical form. in proceedings of the 22nd annual boston university conference on language development (bucld), vol. 22, 97–108. degen, judith & michael k. tanenhaus. 2015. processing scalar implicature: a constraint-based approach. cognitive science 39. 667–710. https://doi.org/10.1111/cogs.12171. fox, danny. 2007. free choice and the theory of scalar implicatures. presupposition and implicature in compositional semantics 71–120. https://doi.org/10.1057/97802302107524. fricke, lea, emilie destruel, malte zimmermann & edgar onea. 2022. the pragmatics of exhaustivity in embedded questions: an experimental comparison of know and predict in german. submitted. geurts, bart. 2010. quantity implicatures. new york: cambridge university press. https://doi.org/10.1017/cbo9780511975158. geurts, bart & nausicaa pouscoulous. 2009. embedded implicatures?!? semantics & pragmatics 2. 1–34. https://doi.org/10.3765/sp.2.4. gronau, quentin f., henrik singmann & eric-jan wagenmakers. 2020. bridgesampling: an r package for estimating normalizing constants. journal of statistical software 92(10). 1–29. https://doi.org/10.18637/jss.v092.i10. kahneman, daniel & amos tversky. 1979. prospect theory: an analysis of decision under risk. econometrica 47(2). 263–291. https://doi.org/10.2307/1914185. klis, martijn van der, bert le bruyn & henriette de swart. 2017. mapping the perfect via translation mining. in proceedings of the 15th conference of the european chapter of the association for computational linguistics, vol. 2, 497–502. https://aclanthology.org/ e17-2080. ‘today at six, einige/#manche people arranged to meet at hasnerplatz.’ proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 59 https://doi.org/10.3765/elm https://www.elm-conference.net/ r core team. 2022. r: a language and environment for statistical computing. r foundation for statistical computing vienna, austria. https://www.r-project.org/. sauerland, uli. 2004. scalar implicatures in complex sentences. linguistics and philosophy 27. 367–391. https://doi.org/10.1023/b:ling.0000023378.71748.db. spector, benjamin. 2007. aspects of the pragmatics of plural morphology: on higher order implicatures. in uli sauerland & penka stateva (eds.), presupposition and implicature in compositional semantics, 243–281. houndsmills et al.: palgrave-macmillan. https://doi.org/10.1057/9780230210752 9. stan development team. 2022a. rstan: the r interface to stan. r package version 2.21.5. https://mc-stan.org/. stan development team. 2022b. stan modeling language users guide and reference manual, version 2.30.1. http://mc-stan.org/. tiel, bob van & bart geurts. 2014. truth and typicality in the interpretation of quantifiers. in proceedings of sinn und bedeutung (sub), vol. 18, 451–468. https://ojs.ub. uni-konstanz.de/sub/index.php/sub/article/view/327. tversky, amos. 1975. a critique of expected utility theory: descriptive and normative considerations. erkenntnis 9(2). 163–173. https://www.jstor.org/stable/20010465. proceedings of elm 2: 49-60, 2023 maya cortez espinoza and lea fricke: on the interpretation of german einige. 60 https://doi.org/10.3765/elm https://www.elm-conference.net/ default biases in the interpretation of english negation, conjunction, and disjunction masoud jasbi, natlia bermudez, & kathryn davidson* abstract. previous research has hypothesized default interpretive biases for three types of ambiguities with english logical words and, or, and not. first, disjunction (a or b) is hypothesized to be biased towardtoward an exclusive interpretation in upward-entailing environments and an inclusive interpretation in downward-entailing environments (levinson 2000, chierchia 2004, breheny et al. 2005). a negated disjunction (not a or b) is claimed to be biased toward a “neither-nor” interpretation (i.e. wide scope negation: ¬[a _ b]) and a negated conjunction is said to be biased toward an “either-not” interpretation (i.e. wide-scope negation: ¬[a ^ b]) (szabolcsi 2002, szabolcsi & haddican 2004). we tested these hypotheses within the same experimental paradigm with 149 english-speaking participants and found disjunction to be biased toward an inclusive interpretation across three different entailment environments: episodic declaratives, questions, and conditional antecedents. our results also confirmed that english negated disjunction is biased toward a “neither-nor” (wide scope negation) interpretation but the results did not show an “either-not” bias (wide scope negation) for negated conjunction. keywords. implicature; disjunction; conjunction; negation; scope; psycholinguistics; semantics; pragmatics 1. introduction. negation, conjunction, and disjunction have played important roles in shaping current theories of semantics and pragmatics. for a long time connective words such as and, or, not and their combinations were considered ambiguous and dissimilar to the semantics of the logical operators in classical logic (tarski 1941). however, grice (1989) presented an alternative theory in which connective words have meanings similar to their logical counterparts, and the apparent ambiguities in their interpretations stem from factors extrinsic to the semantics of the words themselves. we discuss three such ambiguities in this study. first, linguistic disjunction is ambiguous between an “inclusive” and “exclusive” interpretation. a disjunctive statement “a or b” is inclusive if it is interpreted as “a or b or both”, and exclusive if it is interpreted as “a or b, not both”. for example “abe is going to drink tea or coffee” is exclusive if it communicates that he is not going to drink both, and inclusive if it communicates that he may drink both. in most current theories of formal semantics and pragmatics, the meaning of or itself is represented by inclusive disjunction (a _ b), and the exclusive interpretation is the result of a pragmatic enrichment called “scalar implicature”. in the neo-gricean approach (horn 1972, gazdar 1980, levinson 2000), the exclusivity implicature follows logically from three assumptions: 1. that the speaker is cooperative, truthful, informative, relevant, and brief (the gricean maxims); 2. that the speaker could have used and instead (the scalar alternative); and 3. that the speaker is opinionated regarding whether the conjunction is true or not *we would like to thank the organizers and reviewers of elm 2. authors: masoud jasbi, university of california davis (jasbi@ucdavis.edu) & natalia bermudez & kathryn davidson, harvard university. data and code for this study is available at: https://github.com/natibermudez/logic-cards proceedings of elm 2: 129-141, 2023 c©2023 masoud jasbi, natlia bermudez and kathryn davidson published by the lsa with permission of the author(s) under a cc by license. 129 https://doi.org/10.3765/elm https://www.elm-conference.net/ (opinionatedness). in upward entailing environments, the sentence with and is more informative than the sentence with or, while in downward entailing environments this is reversed, so entailment environments are predicted to play a role in whether or receives an exclusive or inclusive reading. although quite different, roughly the same prediction is made in the grammatical approach (chierchia 2004, chierchia et al. 2012, chierchia 2013), where the enrichment is the result of an exhaustivity operator that can appear in various syntactic positions in a sentence’s logical form (may 1985, fox 2008). here too, there is an expectation that structural factors will play a role in interpretation: “the claim is that there are situations in which (standard) implicatures are by default present and situations in which they are by default absent, and such situations are defined by structural factors. by default interpretation, i simply mean the one that most people would give in circumstances in which the context is unbiased one way or the other” (chierchia 2004). entailment environment is one such structural factor: scalar implicatures (e.g. the exclusivity implicature) are present by default in upward entailing environments, but they are suspended in downward entailing environments or environments that license npis such as antecedent of conditions, questions, and the restriction of every. this approach predicts that the default interpretation of disjunction is “exclusive” in upward entailing environments and “inclusive” in downward entailing environments. second, when negation (¬) and disjunction (_) co-occur in a sentence like “abe doesn’t drink tea or coffee”, the sentence can be interpreted in two ways. the first interpretation is a negative disjunction (¬[a _ b]) which we call “the neither-nor interpretation”: “abe drinks neither tea nor coffee”. the second interpretation is the disjunction of negatives ([¬a _ ¬b]) which we call “the either-not interpretation”: “abe either doesn’t drink tea or doesn’t drink coffee; i don’t know which.” szabolcsi (2002) analyzed this ambiguity in terms of scope assignment, and argued that different languages have different default interpretive biases regarding the scope of negation and disjunction. she suggested that in some languages such as hungarian, russian, serbo-croatian, slovak, polish, italian, and japanese the default interpretation is “either-not”. more specifically, in these languages the disjunction words (e.g. the hungarian vagy) are positive polarity items (ppi) and tend to be interpreted outside the scope of negation. in other languages such as english, greek, romanian, bulgarian, and korean, the disjunction words are not ppis and tend to be interpreted in the scope of negation. therefore, the default interpretation in these languages is “neither-nor”. third, when negation (¬) and conjunction (^) co-occur in a sentence like “abe doesn’t drink tea and coffee”, they can also be interpreted in two ways. the first interpretation is negative conjunction (¬[a ^ b]) which is similar to when both is used with and: “abe doesn’t drink both tea and coffee. he drinks one or the other.” this interpretation is essentially the “either-not” interpretation discussed before: “abe either doesn’t drink tea or doesn’t drink coffee, or both”. the second interpretation is the conjunction of negatives [¬a ^ ¬b]) which is equivalent to the “neither-nor” interpretation: “abe doesn’t drink tea and coffee; he drinks neither.” szabolcsi & haddican (2004) proposed that “english disjunction and conjunction happily scope below a c-commanding negation and dutifully obey the de morgan laws, whereas the hungarian counterparts either must scope above the c-commanding negation or fail to obey the de morgan laws. such contrasts are not restricted to english and hungarian. similar to english is german; similar to hungarian are russian, serbian, italian, and japanese, among other languages”. more specifically, they argued that with sentences, quantifiers, and predicates, negated conjunction has a default “either-not” bias. proceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 130 https://doi.org/10.3765/elm https://www.elm-conference.net/ the claims regarding the three ambiguities and their default interpretive biases discussed above are in nature probabilistic. it is easy to show that in each case, both interpretations exist in some context in a given language as examples (1-3) show for english with (a) and (b) continuations. at issue is whether in neutral contexts (i.e. contexts that do not bias the interpretation one way or another), we can find a general bias toward one interpretation rather than another. therefore, while informal and intuition-based judgments can be helpful in forming working hypotheses, a definitive assessment of such probabilistic and empirical claims requires careful psycholinguistic experiments. the next section summarizes prior experimental research on these ambiguities. (1) abe drinks tea or coffee. a. but not both. [a _ b] b. and sometimes both. [a � b] (2) abe doesn’t drink tea and coffee. a. he hates them both. [¬a ^ ¬b] b. he likes only one of them, i forget which one. ¬[a ^ b] (3) abe doesn’t drink tea or coffee. a. he hates them both. ¬[a _ b] b. he likes only one of them, i forget which one. [¬a _ ¬b] 2. previous research. previous experiments on the inclusive vs. exclusive bias for positive disjunction have had mixed results. paris (1973) tested the comprehension of disjunction and other linguistic connectives in children and college students using truth value judgments. in the adult sample, cases of disjunction with or were interpreted as inclusive about 75% percent of the time and those with either-or were inclusive about 67.5% of the time. in two experiments each with 24 participants, evans & newstead (1980) found that “either p or q” type rules are more likely inclusive (52% of the time in experiment 1 and 57% of the time in experiment 2). on the other hand, braine & rumain (1981) used both a give-object task and a truth-value-judgment task and found that adults interpreted disjunction as exclusive most of the time in both experiments (91% and 73% exclusive in the give-object task depending on the phrasing and 41% exclusive in the truth-valuejudgment task). noveck et al. (2002) suggested that lower levels of exclusive interpretations in earlier studies were due to higher task complexity. they collected truth-value judgments from 20 french-speaking adults on abstract logical arguments such as: “if there is a p then there is a q and an r; there is a p, therefore there is a q or an r”. they reported that the majority of the participants (75%) rejected such conclusions because they interpreted the disjunction as exclusive. however, as the authors pointed out, the paradigm was not neutral and encouraged exclusivity implicatures by highlighting the conjunction in the premise of the linguistic stimuli. chevallier et al. (2008) used a truth-value-judgment task with 59 french-speaking participants. testing constructions like “there is an a or a b”, they found that the majority of the responses (about 75%) were inclusive. davidson (2013) used a felicity-judgment task in which 12 english-speakers saw an image and were presented with a sentence like “a spoon is in the mug or a spoon is in the bowl.” participants responded by choosing a smily face or a frowny face to indicate (dis)satisfaction with the linguistic description. the majority of responses (about 80%) proceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 131 https://doi.org/10.3765/elm https://www.elm-conference.net/ showed dissatisfaction with a disjunction when both disjuncts were true which suggests an exclusive interpretation. however, it is also possible that part of the dissatisfaction was due to ignorance implicatures: an utterance of disjunction is odd when the speaker and the addressee already know what is true. jasbi & frank (2021) tested children and adults’ comprehension of disjunction in similar existential sentences to chevallier et al. (2008) using binary and ternary forced-choice truth-value judgments. in the binary task (wrong vs. right) the majority of participants provided an inclusive interpretation (about 80%) but in the ternary task participants chose the intermediate option of “kinda right” more often, which could be interpreted as exclusive. jasbi et al. (2019) varied the number of response options and replicated these findings showing that different rates of exclusivity implicatures can also be due to different number of response options in the truthvalue-judgement task. previous experiments have also reported that disjunction tends to be more inclusive in downward-entailing environments (noveck et al. 2002, schwarz et al. 2007, chemla & spector 2011). regarding the scope of negation and disjunction, lungu et al. (2021) studied four languages experimentally and reported that they did not support the ppi parameter hypothesis of disjunction (szabolcsi 2002, szabolcsi & haddican 2004). they tested 30 french speakers, 43 italian speakers, 27 romanian speakers and 37 english speakers. participants read sentences with simple negative disjunction (e.g. “mary didn’t invite john or suzi”) and had to judge their naturalness on a 7-point likert scale. each sentence had a narrow scope and a wide scope continuation (e.g. “she’s upset with both of them” vs. “i don’t know which of them” respectively). in addition to negation, they tested the effect of other anti-additive operators such as without, few, doubt, and rarely. they found that in all four languages, disjunction exhibits ppi behavior to a certain degree. they concluded that the narrow/wide scope of disjunction is never ruled out completely in a language, and languages differ with respect to the degree to which a particular scope is dispreferred. 3. current study. this study builds on previous research and improves on it in multiple ways. first, it is the first study to test the default biases in the interpretation of negation, conjunction, and disjunction within the same experimental paradigm. second, it uses a card selection task that avoids metalinguistic truth-value judgments and their issues. third, the paradigm avoids linguistic scales (true vs. false, right vs. wrong) that are language dependent and potentially problematic for a cross-linguistic investigation. fourth, it avoids forced-choice responses and allows for a wider range of responses and interpretations. fifth, it also tests the effect of downward-entailing environments on default interpretive biases for negation, conjunction, and disjunction. sixth, the paradigm is extremely simple and does not a-priori bias interpretations one way or another. 3.1. methods. the experiment was designed as a card selection task. participants saw the same six cards to choose from throughout the experiment (figure 1). in each trial, they were presented with a different sentence and were asked to select the cards that best matched it. 3.1.1. linguistic stimuli. we used the words and, or, and not to create 14 logical constructions in english. seven constructions were experimental (table 1) and seven constituted control trial-types (table 2). all constructions used have as the main verb and combined it with one or two of the following nouns: cat, dog, or elephant. for experimental trials, we first considered simple positive and negative constructions. these are constructions like has a cat (positive) or doesn’t proceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 132 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: participants read a sentence prompt (e.g. which has a cat or a dog?) and selected among the cards above. experimental trial-types linguistic (logical) construction simple positive has a [noun] simple negative doesn’t have a [noun] simple positive disjunction has a [noun1] or a [noun2] complex positive disjunction has either a [noun1] or a [noun2] simple negative disjunction doesn’t have a [noun1] or a [noun2] complex negative disjunction doesn’t have either a [noun1] or a [noun2] simple negative conjunction doesn’t have a [noun1] and [noun2] table 1: english logical constructions used in the experimental trials. “noun” was selected from the set of “cat”, “dog”, or “elephant”. have a cat (negative), which compose negation with a simple clause. they help us understand to what extent speakers would interpret the clause exhaustively (e.g. “has only a cat”) vs. nonexhaustively (e.g. “has a cat and possibly some other animal”), and what the effect of negation would be on their meanings and exhaustification. we also considered two types of positive disjunction constructions. one in which the nouns are conjoined only by or (e.g. has a cat or a dog), and one in which they are conjoined by either as well as or (e.g. has either a cat or dog). we call the first construction simple positive disjunction and the second complex positive disjunction. we included the negative variants of both simple and complex disjunction by simply negating the main verb have to doesn’t have. this way we added two negative disjunction constructions: simple negative disjunction (e.g. doesn’t have a cat or a dog) and complex negative disjunction (e.g. doesn’t have either a cat or a dog). finally, we included the simple negative conjunction construction in the experimental trials. simple negative conjunction was similar to simple negative disjunction, except that instead of the disjunction word or we used the conjunction word and (e.g. doesn’t have a cat and a dog). for control trials, we first added the word only to the simple positive trials (e.g. has only a cat). we call these trials “exhaustive simple positive” and use them as controls for “simple positive” trials to capture exhaustive interpretations in our experimental paradigm. we also included their negative versions for the sake of completeness (e.g. doesn’t have only a cat) and called them “exhaustive simple negative” trials. we created “inclusive disjunction” trials by adding the phrase or both to simple disjunction trials (e.g. has a cat or a dog or both). exclusive disjunction trials were created by adding not both to the simple disjunction construction (e.g. has a dog or an eleproceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 133 https://doi.org/10.3765/elm https://www.elm-conference.net/ control trial-types linguistic (logical) construction exhaustive simple positive has only a [noun] exhaustive simple negative doesn’t have only a [noun] inclusive disjunction has a [noun1] or [noun2] or both exclusive disjunction has a [noun1] or [noun2] not both complex negative has neither [noun1] nor a [noun2] simple positive conjunction has a [noun1] and [noun2] complex negative conjunction doesn’t have both a [noun1] and a [noun2] table 2: english logical constructions used in control trials. “noun” was selected from the set “cat”, “dog”, and “elephant”. entailment environment construction example question which ... ... has a cat or a dog? positive episodic bob selected the card(s) which ... ... had a cat or a dog conditional antecedent select a card if ... ... it has a cat or a dog table 3: the three entailment environments tested in this study. phant, not both). these control trials were used to determine when a simple or complex disjunction trial was interpreted as inclusive or exclusive. we created the “complex negative” and “complex negative conjunction” trial-types as controls for negative disjunction and negative conjunction trial-types respectively. the complex negative trial-type used the connective “neither-nor” (e.g. has neither a cat nor an elephant). this construction unambiguously selects for a reading in which both propositions are negated by a connective. complex negative conjunction trials used the word both in addition to negation and conjunction (e.g. doesn’t have both a dog and an elephant). we hypothesized that this construction selects the (wide scope) negation of conjunction in english (¬[a^b]). we used simple conjunction (e.g. has a cat and a dog) as a control trial as well. finally, the control and experimental constructions were embedded in three types of entailment environments: questions, conditional antecedents, and positive episodic statements (table 3). questions were formed using the question word which, for example which has a cat or a dog?. positive episodic statements were reports of actions taken by an imaginary character named “bob”, for example bob selected the card(s) which had a cat or a dog) and participants were asked to copy what the character (i.e. bob) did. with conditionals, the logical construction was in the antecedent and the consequent had the phrase select a card if ..., for example select a card if it has a cat or a dog. this allowed us to experimentally manipulate the entailment environment of these connectives and observe potential effects on the interpretation of the connectives. 3.1.2. participants. we recruited 149 participants online via prolific.co. the study had a between-subjects design for entailment environments; 50 participants saw the linguistic constructions in questions, 50 in conditional antecedents and 49 in positive episodic declaratives (1 participant was excluded from the declarative environment for not finishing the task). in each entailment environment, the study had a within-subjects design and all 50 or 49 participants saw all experiproceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 134 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2: the results for the control (left) and experimental (right) simple positive and negative trials. the x-axis shows the proportion of responses and each color fill a selection of cards. the y-axis shows different trial-types with an example utterance. mental and control constructions. 95 participants reported that they had no prior training in logic. 101 participants were bilingual, 22 monolingual, and 25 multilingual. 3.1.3. procedure. the study had four blocks of questions: the practice block, the experimental block, the control block, and the debriefing block. at first, we asked participants two practice questions. for example in the question environment we asked: 1. which cards have only one animal?, and 2. which cards have two animals? participants were asked to select among the six cards shown in figure 1. these trials helped participants understand the task and know that they can select multiple cards. next, we presented 21 experimental trials randomly (7 trial-types of table 1, 3 trials per trial-type). after the experimental block, participants saw 7 control trials (7 trial-types of table 2, 1 trial per trial-type). the control block followed the experimental block to avoid biasing participant responses. seeing control trials like “cat or dog, or/not both” could potentially bias participants to think more consciously about the inclusive and exclusive interpretations of disjunction. finally, we asked participants whether they were mono/bi/multi-lingual and whether they had any prior training in formal logic. we hypothesized that multilingualism or prior training in logic could affect participant responses. 3.2. results. figure 2 shows the results for the control (left) and experimental (right) simple positive and negative trials in the interrogative, conditional antecedent, and (positive episodic) declarative environments. the left panels show positive exhaustive (e.g. has only a cat) and negative exhaustive control trials (e.g. doesn’t have only a cat). the right panels show simple positive proceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 135 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 3: results for the control (left) and experimental (right) simple and complex positive disjunction trials. the x-axis shows the proportion of responses and each color fill a selection of cards. the y-axis shows different trial-types with an example utterance. (e.g. has a cat) and simple negative (doesn’t have a cat) experimental trials. the results for experimental trials do not match exhaustive controls and instead reflect the basic semantics of simple positive and negative statements without pragmatic enrichment or exhaustification. negative exhaustive trials (e.g. doesn’t have only cat) were (at least) two way ambiguous in all environments: 1. presupposing the prejacent (e.g. has a cat) and asserting the existence of another animal (e.g. dog or elephant); and 2. not presupposing the prejacent (e.g. selecting all cards except the one with a single cat). the presuppositional interpretation was more prevalent in the (positive episodic) declarative environment than the question or the conditional antecedent environments. the presuppositional interpretation was more prevalent in the declarative environment but it was not fully cancelled in the question and conditional antecedent environments either. figure 3 shows the results for the positive simple and complex disjunction trials for the interrogative, conditional antecedent, and declarative environments. the left panels show the inclusive (or both) and exclusive (not both) control trials and the right panels show the experimental trials with simple (e.g. has a cat or a dog) and complex (e.g. has either a cat or a dog) trials. both the simple and complex instances of disjunction were more often (above 75%) inclusive (orange color) than exclusive (green color), but not completely inclusive or exclusive. we used a bayesian mixed-effects logit model with random intercepts and slopes for participants and the fixed effects of disjunction type (simple, complex, inclusive control, exclusive control), entailment environment (question, conditional antecedent, and declarative), and their interaction. the model predicted whether participants included the card with both animals in their selection (i.e. proceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 136 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 4: results for the control (left) and experimental (right) simple and complex negative disjunction trials. the x-axis is the proportion of responses and each color fill a selection of cards. the y-axis shows different trial-types with an example utterance. had an inclusive interpretation)1. the model estimated a positive coefficient for inclusive control disjunction trials compared to simple disjunction trials suggesting that simple disjunction was not completely inclusive (� = 12.00, 95%ci = [3.60, 24.31]) however, a negative coefficient was estimated for exclusive control disjunction compared to simple disjunction suggesting that simple disjunction was not completely exclusive either (� = �15.56, 95%ci = [�27.45,�7.93]). the rate of exclusive interpretations did not differ between simple and complex disjunction or between the different entailment environments (positive episodic declarative vs. question vs. conditionals). in all these cases the 95% credible intervals for the relevant coefficients contained zero. figure 4 shows the results for the negative simple and complex disjunction trials in each linguistic environment. the left panels show the control “neither-nor” (e.g. has neither a cat nor a dog) and “either-not” (e.g. doesn’t have both a cat and a dog) interpretations. the right panel shows the experimental trials for simple negative disjunction (e.g. doesn’t have a cat or a dog) and complex negative disjunction (e.g. doesn’t have either a cat or a dog). the vast majority of responses in both experimental trials received a “neither-nor” interpretation in all three linguistic environments (blue color). we ran a similar bayesian mixed-effects model as in positive disjunction trials, this time predicting the participants’ choice of the “neither-nor” interpretation over other interpretations. the model did not find a difference between negative simple disjunction, negative complex disjunction, and the control “neither-nor” trial-type, since the 95% cred1all our bayesian models used 4 chains, each with 4000 iterations and 2000 as warm-up. for all reported models the chains converged and it was the case that r̂ = 1. proceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 137 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 5: results for the control (left) and experimental (right) negative conjunction trial-type. the x-axis is the proportion of responses, each color fill a selection of cards. the y-axis shows different trial-types with an example utterance. ible intervals for the relevant coefficients contained zero. however, for the “either-not” control trial-type (e.g. doesn’t have both a cat and a dog) the model estimated a negative coefficient, suggesting that participants chose the “neither-nor” interpretation less often in these control trials (� = �17.60, 95%ci = [�27.02,�11.34]). finally, figure 5 shows the results for positive and negative conjunction trials in each linguistic environment. the left panels show the control “neither-nor” and “either-not” interpretations again, as well as the positive conjunction trials (e.g. has a cat and a dog). the right panels show the negative conjunction (e.g. doesn’t have a cat and a dog) trials. the negative conjunction trials show a clear pattern of ambiguity in all three linguistic environments between the “neithernor” (blue color) and the “either-not” interpretations (beige color). we ran a similar bayesian mixed-effects model as before to predict participant choices of “neither-nor” interpretations for negative conjunction trials. the model estimated a positive coefficient for the “neither-nor” control trial-type (� = 13.01, 95%ci = [7.17, 21.54]) and a negative coefficient for the “either-not” control trial-type (� = �8.20, 95%ci = [�19.10,�2.51]). this suggests that participant interpretations of negative conjunction was somewhere between the “neither-nor” and the “either-not” interpretations. the proportion of these interpretations did not vary between different linguistic environments, since the 95% credible intervals for the relevant coefficients contained zero. as expected, participants provided consistent responses to the the positive conjunction trials by choosing the card with both mentioned animals (e.g. cat and dog). we also tested possible effects of mulproceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 138 https://doi.org/10.3765/elm https://www.elm-conference.net/ tilingualism and prior logical training on participant judgments and did not find any effect on the comparisons of interest reported above. 4. discussion. this study tested three hypotheses regarding the proposed default biases in the interpretation of the english logical words and, or, and not. first, a disjunction can be interpreted inclusively or exclusively, and most theories of scalar implicature have suggested that disjunction is biased toward the exclusive interpretation in upward-entailing environments and the inclusive interpretation in downward-entailing environments (levinson 2000, chierchia 2004, breheny et al. 2005). however, this study found that participant interpretations of simple and complex disjunction were biased toward an inclusive interpretation, regardless of the entailment environment. second, a negated disjunction can be interpreted with the negation scoping above the disjunction (¬[a _ b]) or vice versa ([¬a _ ¬b]). previous research had hypothesized that english is biased toward wide scope negation with disjunction (¬[a _ b]) or the “neither-nor” interpretation (szabolcsi 2002). our study confirmed this hypothesis by finding a strong interpretive bias toward the “neither-nor” interpretation in all entailment environments. third, a negated conjunction can be interpreted with negation scoping above the conjunction (¬[a^b]) or vice versa ([¬a^¬b]), and previous research had suggested that in english wide scope negation or the “either-not” interpretation is preferred (szabolcsi & haddican 2004). the results, however, did not show such a preference and the responses were split between the two interpretations. the entailment environment did not affect the scope of negation with conjunction and disjunction either. why did entailment environment not affect the proportion of exclusivity implicatures as predicted? one possibility is that the task failed to properly manipulate the entailment environment and participants largely ignored it. however, such an explanation predicts that there should be no effect of the entailment environment in this study. this is not the case. the rate of presuppositional interpretations for negative exhaustive trials are affected. the exhaustivity of the disjuncts in the control trials of disjunction is also affected. therefore, participants did not ignore the entailment environment altogether. a second possibility is that assuming an independent and default effect of entailment environment on pragmatic implicatures is too strong. entailment environments can and do affect pragmatic implicatures but if only certain other pragmatic or contextual factors are also met. it is possible that different aspects of our experimental setup such as the choice of the main verb have or the task of selecting a group of cards itself has made the entailment environment a less relevant factor for the purpose of implicature computation. finally, this study does not take into account the role of prosody and intonation. it is possible to resolve the ambiguities discussed here using special intonation or at least bias the interpretation one way or another. for the design of this study we had deliberately set aside the issue of prosody to investigate it separately, and after we have a good understanding of how different logical constructions in our study are interpreted when prosody is not explicitly manipulated. therefore, prosody remains a possible confound that we intend to address in future studies. we also plan to expand this paradigm to languages other than english such as hungarian and mandarin chinese with rich literature on logical words. proceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 139 https://doi.org/10.3765/elm https://www.elm-conference.net/ references braine, martin & barbara rumain. 1981. development of comprehension of “or”: evidence for a sequence of competencies. journal of experimental child psychology 31(1). 46–70. breheny, richard, napoleon katsos & john williams. 2005. interaction of structural and contextual constraints during the on-line generation of scalar inferences. in proceedings of the annual meeting of the cognitive science society, vol. 27 27, . chemla, emmanuel & benjamin spector. 2011. experimental evidence for embedded scalar implicatures. journal of semantics 28(3). 359–400. chevallier, coralie, ira noveck, tatjana nazir, lewis bott, valentina lanzetti & dan sperber. 2008. making disjunctions exclusive. quarterly journal of experimental psychology 61(11). 1741–1760. chierchia, gennaro. 2004. scalar implicatures, polarity phenomena, and the syntax/pragmatics interface. structures and beyond 3. 39–103. chierchia, gennaro. 2013. logic in grammar: polarity, free choice, and intervention. oup. chierchia, gennaro, danny fox & benjamin spector. 2012. scalar implicature as a grammatical phenomenon. in handbücher zur sprach-und kommunikationswissenschaft/handbooks of linguistics and communication science semantics volume 3, de gruyter. davidson, kathryn. 2013. and or or: general use coordination in asl. semantics & pragmatics . evans, jonathan & stephen newstead. 1980. a study of disjunctive reasoning. psychological research 41(4). 373–388. fox, danny. 2008. on logical form. in randall hendrick (ed.), minimalist syntax, 82–123. wiley-blackwell. gazdar, gerald. 1980. pragmatics and logical form. journal of pragmatics 4(1). 1–13. grice, paul. 1989. studies in the way of words. harvard university press. horn, laurence r. 1972. on the semantic properties of logical operators in english: ucla dissertation. jasbi, masoud & michael c frank. 2021. adults’ and children’s comprehension of linguistic disjunction. collabra: psychology 7(1). 27702. jasbi, masoud, brandon waldon & judith degen. 2019. linking hypothesis and number of response options modulate inferred scalar implicature rate. frontiers in psychology 10. 189. levinson, stephen c. 2000. presumptive meanings: the theory of generalized conversational implicature. mit press. lungu, oana, anamaria fălăus, & francesca panzeri. 2021. disjunction in negative contexts: a cross-linguistic experimental study. journal of semantics 38(2). 221–247. may, robert. 1985. logical form: its structure and derivation. mit press. noveck, ira, gennaro chierchia, florelle chevaux, raphaëlle guelminger & emmanuel sylvestre. 2002. linguistic-pragmatic factors in interpreting disjunctions. thinking & reasoning 8(4). 297–326. paris, scott g. 1973. comprehension of language connectives and propositional logical relationships. journal of experimental child psychology 16(2). 278–291. schwarz, florian, charles clifton jr & lyn frazier. 2007. strengthening’or’: effects of focus and downward entailing contexts on scalar implicatures. university of massachusetts occasional proceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 140 https://doi.org/10.3765/elm https://www.elm-conference.net/ papers in linguistics 33(1). 9. szabolcsi, anna. 2002. hungarian disjunctions and positive polarity. in istvan kenesei & peter siptar (eds.), approaches to hungarian, vol. 8, university of szeged. szabolcsi, anna & bill haddican. 2004. conjunction meets negation: a study in cross-linguistic variation. journal of semantics 21(3). 219–249. tarski, alfred. 1941. introduction to logic and to the methodology of the deductive sciences. oup. proceedings of elm 2: 129-141, 2023 masoud jasbi, natlia bermudez and kathryn davidson: default biases in the interpretation of english negation, conjunction, and disjunction. 141 https://doi.org/10.3765/elm https://www.elm-conference.net/ fake reefs are sometimes reefs and sometimes not, but are always compositional hayley ross, najoung kim & kathryn davidson* abstract. the semantics of adjective modification often begins with set intersection, such that jyellow flowerk = jyellowk∩jflowerk. thus a yellow flower is a flower. such an account, however, runs into problems for adjectives like fake or counterfeit, which display a privative inference: a fake gun is not a gun and a counterfeit dollar is not a dollar. moreover, recent work shows privativity cannot easily be encoded as a property of specific adjectives like counterfeit, since e.g. counterfeit watch robustly licenses the subsective inference of being a watch (martin 2022). we gather judgments on nearly 800 adjective-noun bigrams (of which 180 are novel, i.e. zero corpus frequency), and show that privativity depends on the adjective, noun and context, and can be manipulated for the very same adjective-noun bigram by presenting it in different contexts. this poses a challenge for theories which fix privativity as a property of the adjective and always use the same method of composition (partee 2010, del pinal 2015). moreover, we find no difference in participant behavior between novel adjective-noun bigrams and high frequency ones, suggesting that the process is nonetheless compositional and not the result of convention or memorized idiosyncrasy. our results support compositional accounts like martin (2022) (which modifies del pinal 2015) and guerrini (2024), which treat privativity as context-dependent. keywords. adjectives; nouns; compositionality; privativity; entailment; semantics 1. introduction. a central concern for the study of meaning is how the meanings of complex expressions are composed from the meanings of their constituent parts. the fact that people understand completely novel phrases provides an argument that meaning must be governed by some kind of compositionality (partee 2009). this paper, following in a growing tradition (partee 2009, 2010, szabó 2012, del pinal 2015, i.a.), studies the dynamic interaction of meaning and context through the lens of (privative) adjective modification and how to account for it compositionally. historically, privativity has been defined as an adjective-specific phenomenon which negates the noun that the adjective combines with. a fake gun is said to be precisely not a gun. this property distinguishes privative adjectives from other types of adjectives, which typically yield an intersective or subsective inference. canonical examples of privative adjectives include fake, false, former, counterfeit, knock-off, mock, and perhaps also artificial and virtual (nayak et al. 2014). (1) intersective inference this is a yellow flower. ∴ this is yellow. ∴ this is a flower. (2) subsective inference this is a small elephant. ∴ this is an elephant. ̸∴ this is small. (3) privative inference this is a fake gun. ∴ this is not a gun. *many thanks to research assistant kate bigley, and to jesse snedeker and all of the members of the harvard meaning & modality lab for their helpful feedback. special thanks to josh martin, whose dissertation on this topic inspired this project. this work was supported by an mbb graduate student research award from harvard’s mind, brain and behavior initiative. authors: hayley ross, harvard university (hayleyross@g.harvard.edu), najoung kim, boston university (najoung@bu.edu) & kathryn davidson, harvard university (kathryndavidson@fas.harvard.edu). proceedings of elm 3: 332-343, 2025 c©2025 hayley ross, najoung kim, and kathryn davidson published by the lsa with permission of the author(s) under a cc by license. 332 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ privativity poses a challenge for compositionality, which requires the meaning of a complex expression to be derived solely from the meaning of its constituent parts (szabó 2012): the set of guns combined with whatever function or set fake denotes. modification of nouns by adjectives is classically treated as simple set intersection, shown in (4), as in textbooks like heim & kratzer (1998) and coppock & champollion (2023).1 if a fake gun is not a gun, it is not clear how to derive the meaning of fake gun from the set of (real) guns through set operations such as subsection. this means that our compositional process cannot arrive at the meaning of fake fire simply by having fake yield some subset of the meaning of fire. (4) jyellow flowerk = jyellowk ∩ jflowerk = {x : x is yellow} ∩ {x : x is a flower} in contrast to the conventional framing of privativity as an adjective-specific phenomenon, martin (2022) shows that inference patterns for so-called privative adjectives vary depending on the noun used. for example, counterfeit may license a privative or subsective inference depending on the noun (and accompanying context). (5) subsective inference this is a counterfeit watch. ∴ this is a watch. (6) privative inference this is a counterfeit dollar. ∴ this is not a dollar. this per-adjective and per-noun variation raises an additional question of whether these adjectivenoun combinations and their inferences are computed (compositionally) on the fly, based on just the given adjective, noun and context, or whether there is an element of convention or past experience necessary to derive these varying inferences, in which case the inference would be stored (memorized). a significant body of processing work (arnon & snider 2010, tremblay & baayen 2010, caldwell-harris et al. 2012, o’donnell 2015, i.a.) reveals plenty of cases where humans don’t appear to compose meaning on the fly: chunks of various sizes from multi-morpheme words to entire idiomatic expressions, especially highly frequent words or expressions, can get stored as units and trigger priming effects in experimental studies, whether their meaning is idiomatic or fully compositional from their parts. if the effect of adjectives with privative inferences is stored rather than composed on the fly, then deriving the inference for infrequent adjective-noun bigrams with such adjectives, such as fake scarf or fake reef, might be difficult or result in widely varying results between people. the same might be true for intermediate, less memorization-heavy approaches such as learning (memorizing) the inferences for some high-frequency bigrams and then reasoning about novel bigrams by analogy where possible.2 this paper explores the effect of experience (as measured by corpus frequency of the bigram) and context on adjective-noun combination and inferences, especially for novel (zero corpus frequency) adjective-noun bigrams. we gather a large quantity of adjective-noun inference judgments for both high-frequency and novel / zero corpus frequency adjective-noun bigrams over three experiments and show that inferences depend not just on the adjective but also on the noun and the context. further, we show that novel adjective-noun bigrams and their privativity inferences are handled as productively and 1to their credit, both textbooks note that gradable and/or non-intersective adjectives are not handled by this account. 2for example, a participant might reason that the novel bigram counterfeit scarf is a scarf by analogy to other clothing items and accessories such as watch or handbag which they have seen counterfeit occur with subsectively. proceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 333 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ consistently by participants as those of high-frequency ones, despite the significant variation by adjective, noun, and context. thus, any theory of adjective-noun combination must address the context-sensitivity of these inferences and predict the ability to generalize to novel combinations, for example using a compositional approach; it cannot be based on memorized idiomatic meanings. in section 5, we discuss two compositional accounts of adjective-noun modification which satisfy these requirements and can handle the new data presented in this paper. 2. choice of adjective-noun bigrams. we first establish a set of 798 bigrams to study, 23% of which are zero frequency in a large corpus, thus presumed novel for participants. the full set of bigrams, as well as results for all experiments in this paper, are available on github.3 2.1. selection by corpus frequency. we consider 6 “privative” adjectives of interest: fake, counterfeit, false, artificial, knockoff and former. since we established in the introduction that such adjectives need not always be privative, from here on out the phrase “privative adjective” will refer to these adjectives such as fake that have been discussed in prior literature as (typically) resulting in privative inferences. we select 6 intersective/subsective adjectives as “controls” which each have a similar frequency to one of the privative adjectives in a very large corpus (c4, ca. 130 trillion words; raffel et al. 2020, dodge et al. 2021) and which have relatively few selectional restrictions: useful, tiny, illegal, homemade, unimportant and multicolored. frequencies are shown in table 1. we choose multicolored as a low-frequency example of a (standardly intersective) colour adjective, illegal since it has negative valency while typically being subsective, and homemade since it targets the manner of manufacture, similar to counterfeit and artificial, without being obviously privative. adjective tokens adjective tokens former 15.8m useful 13.6m false 4.6m tiny 5.8m artificial 3.9m illegal 4.5m fake 3.1m homemade 2.2m counterfeit 450k unimportant 170k knockoff 57k multicolored 93k table 1: adjective frequencies in the c4 corpus (130t words) we algorithmically select 43 nouns from 300 nouns which commonly occur with a wide range of adjectives (pavlick & callison-burch 2016a), plus the 36 nouns used in martin (2022), with the goal of generating a high number of zero-frequency bigrams. we then manually select an additional 59 nouns which are semantically similar to these 43 nouns, for a total of 102 nouns. we cross these 102 nouns with the 12 adjectives for a total of 1224 bigrams to use in subsequent experiments. we determine (relative) bigram frequency by counting the frequency of all bigrams involved in this process in c4, for a total of 3979 bigrams (358 unique nouns × 12 adjectives, plus experiment fillers), since calculating the frequency over every possible corpus bigram would be prohibitively expensive. thus, terms like “high frequency bigram” or “top quartile bigram” in this paper should be interpreted in relative rather than absolute terms. 3https://github.com/rossh2/artificial-intelligence proceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 334 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 2.2. experiment 1. experiment 1 filters out clearly nonsensical combinations resulting from the blind crossing of adjectives and nouns. combinations like counterfeit accusation or multicolored effort must be excluded before we can reasonably ask questions like “is a counterfeit n still an n?” participants are presented with a bigram and asked “how easy is it to imagine what this would mean?”, as shown in figure 1a.4 144 native american english speakers5 were recruited on prolific (of which 7 were excluded due to failed attention checks and/or not meeting the criteria for native english speaker); the study was implemented in qualtrics. participants were paid pro rata at $12/hour; the experiment took 4 minutes on average. each participant saw 14 questions (12 target bigrams, 2 fillers), for 3 ratings per bigram in total. we categorize bigrams whose ratings were majority “very hard” or “somewhat hard” as nonsensical, and exclude them from subsequent experiments. this leaves 798 bigrams, of which 23% (180) are zero frequency in c4, i.e. almost certainly novel to new participants, and another 21% (170) are low frequency (bottom quartile), so also quite possibly novel to participants. (a) experiment 1 (b) experiment 2 on pcibex figure 1: screenshots of questions in experiments 1 and 2 3. experiment 2: is an a n an n?. 3.1. method. experiment 2 asks participants is an a n still an n? for each of the 798 adjectivenoun bigrams left after filtering in experiment 1. an example question is shown in figure 1b. we choose to use the same design as martin (2022), with the slight modification of adding still to make the question more natural.6 this differs from previous privativity studies (pustejovsky 2013, pavlick & callison-burch 2016a,b) which ask participants to rate these inferences given a particular sentence context drawn from a corpus. for example, pavlick & callison-burch (2016b) ask whether (7-a) entails (7-b). while more realistic than out-of-the-blue judgments, this also creates a much noisier picture, as demonstrated by this example: participants rate (7-a) as contradicting (7-b), but the verb denied and the world knowledge of pharmacists selling medicine also seem to be driving part of this inference. it actually remains unclear whether the counterfeit medicine that the pharmacists were selling qualifies as medicine in this scenario. 4previous work studying novel adjective-noun combinations (vecchi et al. 2017) uses a more complex pairwise ranking approach to precisely measure semantic deviance, but we only need to filter out obviously nonsensical bigrams. 5we recruit people on prolific who self-report english as their first and primary language and are located in the united states. we further ask them at the end of the study whether they learned english before the age of 5 and whether they speak american english as opposed to another dialect of english (if not, they are paid but excluded). this implementation of “native speaker” is merely intended as a practical way to expect shared language experiences among our participant sample (cheng et al. 2021). 6using still seems to help foreground the idea that adding the adjective might change noun membership: a fake scarf or unimportant sign might not be a scarf or sign, or conversely might be a scarf despite being fake. proceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 335 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (7) a. pharmacists in algodones denied selling counterfeit medicine in their stores. b. pharmacists in algodones denied selling medicine in their stores. this experiment aims to show, for a wider range of adjectives and nouns than martin 2022, that privativity varies depending on the noun, and investigates whether participants behave differently for high and zero frequency (assumed to be novel) bigrams. we will study non-out-the-blue judgments (in a more controlled setting than pavlick & callison-burch) later, in experiment 3. we ran this experiment in two parts. for the first 305 bigrams, 510 native american english speakers were recruited on prolific (of which 15 were excluded due to failed attention checks and/or not meeting the criteria for native english speaker); the study was implemented in pcibex (zehr & schwarz 2018). for the next 498 bigrams, 756 native american english speakers were recruited on prolific (of which 24 were excluded); this study was implemented on qualtrics. participants were paid pro rata at $12/hour and the experiment took 3 minutes on average. each participant saw 12 questions (4 typically-intersective adjectives, 4 typically-privative adjectives, 4 fillers), for a total of approx. 12 ratings/bigram.7 since some bigrams which may not make sense to everyone likely remain after experiment 1, we explicitly alert participants in experiment 2 to this possibility and instruct them to use the “unsure” rating if a combination does not make sense to them. 3.2. results. mean bigram ratings are shown in figure 2 (organized by adjective, each dot represents the mean rating for one adjective-noun bigram), and individual ratings by participants for a selection of bigrams are shown in figure 3. we find that each so-called “privative” adjective in fact yields graded variation from privative to subsective depending on the noun, with ratings spanning all the way from 1 (“definitely not [an n]”) to 5 (“definitely yes [an n]”). in figure 3, we can also see that intermediate means are often associated with high variance rather than participants agreeing on “unsure”. we also find that “subsective” adjectives are usually subsective (“probably yes” or “definitely yes”), warranting the name, but are nonetheless not so clearly subsective with certain nouns (e.g. homemade cat with µ = 2.6, illegal currency with µ = 2.83). secondly, we find no effect of bigram frequency on rating variance. a linear regression in r (r core team 2023) shows that bigram frequency correlates poorly with the variance in the ratings (typically subsective: r2 = 0.003, typically privative: r2 = 0.010, both p > 0.05). instead, participants agree to a similar degree on the meaning and inferences for high-frequency and zero frequency (novel) adjective-noun bigrams. some high frequency bigrams such as artificial tree or former house show high variance in ratings (µ = 3.50, σ2 = 1.83 and µ = 3.63, σ2 = 1.76 respectively), suggesting that these bigrams do not have a conventionalized meaning or inference when presented out of the blue. moreover, some zero frequency bigrams like knockoff image and counterfeit scarf have quite low variance (µ = 4.90, σ2 = 0.10 and µ = 4.80, σ2 = 0.18), suggesting that participants compose even novel bigrams systematically. 3.3. discussion. the results from experiment 2 lend further weight to previous work illustrating that no adjective is unequivocally privative, but rather that privativity depends on the combination of adjective and noun (martin 2022). we further see no correlation between rating variance and bigram frequency. we conclude that high frequency need not lead to a fixed conception of bi7due to issues with pcibex, the first part of experiment did not yield an even number of ratings per bigram. in the analysis of this experiment, we randomly sample and cap the number of ratings at (10-)12 ratings/item. proceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 336 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 2: mean ratings for “is an an an n?” for each bigram by adjective in experiment 2, where 1 is most privative (“definitely not”) and 5 is most subsective (“definitely yes”). figure 3: participant ratings for a selection of bigrams involving fake and counterfeit in experiment 2, where 1 is most privative and 5 is most subsective. • shows the mean with se for each bigram. gram meaning or inference, and that previous exposure to a (potentially) privative adjective-noun pair is not required to draw this inference, even for adjectives with relatively broad meanings like fake. instead, we suspect that high variance may be due in part to participants imagining different contexts for the bigrams (which were presented out of the blue in experiment 2), such that e.g. fake might target different aspects of the noun’s properties or different properties might be relevant for noun-hood in that context. for example, a fake crowd might qualify as a crowd if it is made up of paid actors, but less so if it is just painted dummies on a movie set. 4. experiment 3: context. 4.1. method. for 28 adjective-noun bigrams from experiment 2, we construct two contexts each intended to bias the reader towards a subsective or privative inference respectively. two example contexts for fake concert are shown in figure 4. we targeted 6 pairs of adjective-noun bigrams from experiment 2 with intermediate mean ratings and high variance, such that one bigram is zero/low frequency and the other is high frequency; we will use these pairs to investigate any effect of frequency. we then selected an additional 16 bigrams with intermediate mean ratings and high variance for which we were able to write convincing example contexts. we select proceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 337 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (a) subsective-biased context (b) privative-biased context figure 4: screenshots of contexts for fake concert in experiment 3. these bigrams which are neither at ceiling or floor precisely because we suspect that there may be more than one context in participants’ minds, and thus the specified contexts might split apart these middle-of-the-scale ratings into high (subsective) or low (privative) ratings respectively, explaining (some of) the variance in experiment 2. further, we select pairs of high and zero/low frequency bigrams because it is possible that high frequency bigrams might come with more conventionalized contexts and/or more fixed meanings and inferences in general, and thus might resist manipulation by provided contexts. conversely, zero/low frequency bigrams might be particularly easy to manipulate, since they lack any preconceived “default” context. for the first 12 bigrams, 40 native american english speakers were recruited on prolific (of which 1 excluded due to failed attention checks); for the second set of 18 bigrams (two bigrams were rerun), a further 40 native american english speakers were recruited on prolific (of which 2 were excluded for not meeting the native speaker criteria). both studies were implemented in qualtrics. participants were paid pro rata at $12/hour; the experiment took 8 minutes on average. in the first instance of the experiment, each participant saw 12 items as shown in figure 4; in the second instance, each participant saw 18 items (3 or 6 intersective-biased, 3 or 6 privative-biased, 6 fillers), yielding 10 ratings/item. 4.2. results. we find that across the board, writing biased contexts does indeed shift participants ratings in the intended direction, as shown in figure 5, as well as reducing the variance. the one exception is counterfeit dollar, which refuses to be influenced by context at all. this can be explained simply by its meaning: dollars depend so heavily on having an authentic method of manufacture that any way in which they can be counterfeited, i.e. in which their method of manufacture is non-conventional, robs them of being a dollar. we fit an ordinal mixed effects model in r (r core team 2023, christensen 2022) and find statistically significant effects for both the subsective-biased and privative-biased contexts compared to having no context (p < 0.05 for both; a subsective-biased context makes a high (subsective) rating 4x more likely while a privative biased context makes a high (subsective) rating only 1⁄6x as likely). as in experiment 1, we find no effect of frequency in this experiment: high-frequency bigrams do not have more fixed inferences and are not more resistant to being manipulated by context. proceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 338 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 5: ratings for “in this setting, is an an still an n?” for 12 of 28 bigrams in experiment 3, where 1 (“definitely not”) is most privative and 5 (“definitely yes”) is most subsective. the ratings from experiment 2 are shown in gray. the first and third columns show zero or low frequency bigrams; the second and fourth columns show corresponding high frequency bigrams. 4.3. discussion. the results from experiment 3 show that the subsective/privative inferences drawn from adjective-noun combination are indeed context-dependent: providing different contexts can cause participants to draw quite subsective or privative inferences for the very same adjective-noun bigram. this is possible whether the bigram is high-frequency (thus potentially coming with a “default” or conventionalized context of use) or novel. thus it is likely that the participants’ imagined contexts explains some of the variance in experiment 2, which presented the bigrams out of the blue. in other words, the meaning of adjective-noun bigrams and their privative inferences cannot be explained by memorization of a single (conventionalized) meaning or inference, since inferences must be computed productively on the fly based on the provided context (as well as world knowledge). finally, the ability to manipulate the inferences of novel bigrams such as false concert supports a compositional account for the meaning of the bigram (where context is included as part of the composition and inference-drawing process), as in experiment 2. 5. impact on theoretical accounts. our experiments showcase the wide variation in privative inferences among so-called privative adjectives: first, adjectives can license either a subsective or privative inference depending on the noun (and context), and second, the same adjective-noun bigram can license either inference depending on context. these results support a compositional, context-dependent account of adjective-noun modification rather than an approach based on prior experience or convention which memorizes the meanings and/or inferences for previously encountered bigrams. further, these results pose challenges for theories which treat privativity as a property of only the adjective, such as partee (2010), del pinal (2015) and guerrini (2024). 5.1. partee’s non-vacuity principle. partee’s classic account of privative adjectives (kamp & partee 1995, partee 2007, 2009, 2010) posits that all seemingly privative adjectives in fact compose subsectively with the noun. unlike regular subsective adjectives, however, jfake gunk = ∅ initially, since gun includes only real guns and fake is privative (by definition). the non-vacuity proceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 339 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ principle then coerces gun to expand to include both real and fake guns, so that fake can now act subsectively over this new expanded set. since this account stipulates that fake is always privative with respect to the original noun, it is unclear how it accounts for our experimental data that e.g. a fake watch is frequently a watch. while the intuitions behind partee’s account seem to be on the right track and have inspired several subsequent analyses, her specific implementation and the non-vacuity principle do not capture the full extent of our data, in addition to existing criticisms of her implementation (martin 2022, del pinal 2015, 2018, guerrini 2024). 5.2. qualia-based accounts. del pinal (2015, 2018) argues that we can have a compositional, truth-conditional account and still explain the behaviour of adjectives like real, typical and fake in a principled manner by moving to a two-dimensional semantics. del pinal’s dual content semantics enriches lexical entries to have an extensional component (e-structure), which corresponds to the traditional set extension, plus a conceptual component (c-structure). this cstructure essentially captures the concept behind the noun or adjective, leaning on the large body of literature on concepts in psychology to do so. del pinal implements it using qualia. adjectives like fake draw on the c-structure of nouns like gun to build the new e-structure for fake gun. specifically, fake modifies gun to yield the semantics in (8) for fake gun: firstly, in the e-structure, fake guns are not in the extension of guns. secondly, fake guns do not have the origins of guns (the agentive property), instead, they were made to appear as if they were guns. in the c-structure, this is also the new agentive property, and means that they have the appearance of guns (the same formal properties) and do not have the same purpose (telic property). (8) jfake gunk = (del pinal 2015; p.21) e-structure: λx.¬qe(jgunk)(x) ∧ ¬qa(jgunk)(x) ∧ ∃e2 [making(e2) ∧ goal(e2, qf (jgunk)(x))] c-structure: constitutive: qc(jgunk) = λx. parts-gun(x) formal: qf (jgunk) = λx. perceptual-gun(x) telic: ¬qt (jgunk) = λx.¬gen e[shooting(e) ∧ instrument(e, x)] agentive: λx.∃e2 [making(e2) ∧ goal(e2, qf (jgunk)(x))] while del pinal’s theory lays out in much more detail than partee how the composition works, del pinal (2015) still stipulates, like partee, that fake guns are not guns by fixing in the e-structure that a fake gun is not in the extension of gun: ¬qe(jgunk)(x). in subsequent work, del pinal (2018) admits that this part of the e-structure is questionable for counterfeit and artificial and that it is an empirical question whether this should be included or not. empirically, both we and martin (2022) find that it should not be included for any “privative” adjective. martin (2022) adjusts del pinal’s definitions to remove privativity in the e-structure. instead, privativity arises after composition of the adjective and noun, when set membership of the newly composed object is determined. if a targeted, negated quale (primarily telic for fake) is “central” to the meaning of the noun, i.e. part of the e-structure, then this results in privativity. which qualia (analogous to “typical properties” of the noun) are incorporated into the e-structure is contextdependent, as it is for any use of a noun. this allows the modified account to capture the variation in our experiments by noun and by context. proceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 340 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 5.3. similarity-based accounts. guerrini (2024) presents a similarity-based account of fake, which builds on the intuitions for fake and counterfeit in del pinal (2015), but implements the lexical entry using using the one-dimensional compositional semantics for seems like from guerrini (2022) instead of using a two-dimensional semantics with qualia. (9) jfakek = λpλx. intended(x, seem-like(x, p )) ∧ ¬p (x) (guerrini 2024; p.190) (10) intended(x, p ) = ∃e. action(e, x) ∧ goal(e, p (x)) note that like the original definition of fake in del pinal (2015), and in keeping with partee, privativity is explicitly part of guerrini’s definition of fake. unlike del pinal, guerrini is committed to this fact. he argues that the variability in privativity in the experiments in martin (2022), which we extend in this paper, is in fact explained by syntactic ambiguity. building on the account in martin (2022) for the ambiguity between subsective and intersective readings of e.g. good thief (good at thieving, or a good person and a thief), guerrini argues that when fake watch has a subsective reading, fake (privatively) modifies another covert, contextually supplied noun, such as rolex: syntactically, we have [np [ap fake rolex] watch]. since this syntactic ambiguity is context-dependent, guerrini (2022) in principle also accounts for the data presented in experiments 2 and 3. one concern with this account is whether the contextually supplied material has to be a single noun over which fake acts privatively. our data suggests this would be difficult for items such as fake concert in figure 4, though there may be a multi-word phrase or concept that can satisfy the theory. 6. conclusion. this paper presents experimental evidence on the variation in subsective vs. privative inferences both within adjectives and within adjective-noun bigrams, including for novel adjective-noun bigrams. we find that no adjective always yields privative inferences, lending further weight to martin (2022). we further find that for most adjective-noun combinations, privativity depends on the context as well as the adjective and the noun. our results show that any theory of adjective-noun combination must account for the context-sensitivity of these inferences and allow generalization to novel combinations (for example, by composition)—these inferences are not so unpredictable as to need to be memorized. theories of privativity which are not compositional (e.g. basic set complementation) or which fix privativity as a property of individual adjectives (partee 2010, del pinal 2015) with only a single method of composition do not account for the full set of our data. compositional accounts like martin (2022)’s modification of del pinal (2015, 2018) and guerrini (2024) which predict that privativity is context-dependent, either by having privativity arise outside of the composition or by appealing to syntactic ambiguity, are able to account for the generalization and context-sensitivity found in our experiments. these compositional approaches aim to explain why participants are equally able to draw inferences for novel bigrams as for highfrequency ones, and what inferences we should expect given the effect of the context on available nouns or restrictions of noun denotations. references arnon, inbal & neal snider. 2010. more than words: frequency effects for multi-word phrases. journal of memory and language 62(1). 67–82. https://www.sciencedirect.com/ science/article/pii/s0749596x09000965. caldwell-harris, catherine, jonathan berant & shimon edelman. 2012. measuring mental enproceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 341 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ trenchment of phrases with perceptual identification, familiarity ratings, and corpus frequency statistics. in dagmar divjak & stefan th. gries (eds.), frequency effects in language representation, vol. 2, 165–194. berlin: de gruyter mouton. https://www. degruyter.com/document/doi/10.1515/9783110274073.165/html. cheng, lauretta s. p., danielle burgess, natasha vernooij, cecilia solı́s-barroso, ashley mcdermott & savithry namboodiripad. 2021. the problematic concept of native speaker in psycholinguistics: replacing vague and harmful terminology with inclusive and accurate measures. frontiers in psychology 12. https://www.frontiersin.org/journals/ psychology/articles/10.3389/fpsyg.2021.715843/full. christensen, rune haubo bojesen. 2022. ordinal—regression models for ordinal data. manuscript. coppock, elizabeth & lucas champollion. 2023. invitation to formal semantics. https:// eecoppock.info/semantics-boot-camp.pdf. dodge, jesse, maarten sap, ana marasović, william agnew, gabriel ilharco, dirk groeneveld, margaret mitchell & matt gardner. 2021. documenting large webtext corpora: a case study on the colossal clean crawled corpus. http://arxiv.org/abs/2104.08758. guerrini, janek. 2022. ’like a n’ constructions: genericity in similarity. in dean mchugh & alexandra mayn (eds.), proceedings of the esslli 2022 student session, galway, ireland. https://doi.org/10.21942/uva.20368104. guerrini, janek. 2024. keeping fake simple. journal of semantics 41(2). 175–210. https: //doi.org/10.1093/jos/ffae010. heim, irene & angelika kratzer. 1998. semantics in generative grammar blackwell textbooks in linguistics 13. malden, mass., usa: blackwell. kamp, hans & barbara h. partee. 1995. prototype theory and compositionality. cognition 57(2). 129–191. martin, joshua. 2022. compositional routes to (non)intersectivity. cambridge, ma: harvard university dissertation. https://www.proquest.com/docview/2681380157/ abstract/58a63b8c3e6548aepq/1. nayak, neha, mark kowarsky, gabor angeli & christopher d. manning. 2014. a dictionary of nonsubsective adjectives. tech. rep. cstr 2014-04 department of computer science, stanford university. https://www-cs.stanford.edu/˜angeli/papers/ 2014-tr-adjectives.pdf. o’donnell, timothy j. 2015. productivity and reuse in language: a theory of linguistic computation and storage. cambridge, massachusetts, london, england: the mit press. https://doi.org/10.7551/mitpress/9780262028844.001.0001. partee, barbara h. 2007. compositionality and coercion in semantics: the dynamics of adjective meaning. in gerlof bouma, irene krämer & joost zwarts (eds.), cognitive foundations of interpretation, 145–161. amsterdam: royal netherlands academy of arts and sciences. partee, barbara h. 2009. formal semantics, lexical semantics, and compositionality: the puzzle of privative adjectives. philologia 7(1). 11–21. http://www.philologia.org.rs/ index.php/ph/article/view/216. partee, barbara h. 2010. privative adjectives: subsective plus coercion. in thomas zimproceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 342 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ mermann, rainer bauerle & uwe reyle (eds.), presuppositions and discourse: essays offered to hans kamp, 273–285. leiden, the netherlands: brill. https://brill.com/ downloadpdf/book/edcoll/9789004253162/b9789004253162-s011.pdf. pavlick, ellie & chris callison-burch. 2016a. most “babies” are “little” and most “problems” are “huge”: compositional entailment in adjective-nouns. in katrin erk & noah a. smith (eds.), proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers), 2164–2173. berlin, germany: association for computational linguistics. https://aclanthology.org/p16-1204. pavlick, ellie & chris callison-burch. 2016b. so-called non-subsective adjectives. in claire gardent, raffaella bernardi & ivan titov (eds.), proceedings of the fifth joint conference on lexical and computational semantics, 114–119. berlin, germany: association for computational linguistics. https://aclanthology.org/s16-2014. del pinal, guillermo. 2015. dual content semantics, privative adjectives, and dynamic compositionality. semantics and pragmatics 8. 7:1–53. https://semprag.org/index.php/ sp/article/view/sp.8.7. del pinal, guillermo. 2018. meaning, modulation, and context: a multidimensional semantics for truth-conditional pragmatics. linguistics and philosophy 41(2). 165–207. https://doi. org/10.1007/s10988-017-9221-z. pustejovsky, james. 2013. inference patterns with intensional adjectives. in harry bunt (ed.), proceedings of the 9th joint iso acl sigsem workshop on interoperable semantic annotation, 85–89. potsdam, germany: association for computational linguistics. https: //aclanthology.org/w13-0509. r core team. 2023. r: a language and environment for statistical computing. https:// www.r-project.org/. raffel, colin, noam shazeer, adam roberts, katherine lee, sharan narang, michael matena, yanqi zhou, wei li & peter j. liu. 2020. exploring the limits of transfer learning with a unified text-to-text transformer. journal of machine learning research 21(140). 1–67. http://jmlr.org/papers/v21/20-074.html. szabó, zoltán gendler. 2012. the case for compositionality. in wolfram hinzen, edouard machery & markus werning (eds.), the oxford handbook of compositionality, 64–80. oxford: oup. https://doi.org/10.1093/oxfordhb/9780199541072.013.0003. tremblay, antoine & harald baayen. 2010. holistic processing of regular four-word sequences: a behavioural and erp study of the effects of structure, frequency, and probability on immediate free recall. perspectives on formulaic language: acquisition and communication 151–173. vecchi, eva m., marco marelli, roberto zamparelli & marco baroni. 2017. spicy adjectives and nominal donkeys: capturing semantic deviance using compositionality in distributional spaces. cognitive science 41(1). 102–136. https://onlinelibrary.wiley.com/ doi/abs/10.1111/cogs.12330. zehr, jeremy & florian schwarz. 2018. penncontroller for internet based experiments (ibex). https://doi.org/10.17605/osf.io/md832. proceedings of elm 3: 332-343, 2025 hayley ross, najoung kim, and kathryn davidson: fake reefs are sometimes reefs and sometimes not, but are always compositional. 343 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ does ‘a couple’ pattern with scalars or numbers insights from inference and ‘so’ tasks erying qin, chao sun & richard breheny* abstract. previous research establishes that paucal quantifiers like ‘a couple’ are ambiguous between the literal meaning of ‘at least two’ and the enriched meaning understood as conveying a restriction on quantity, the latter of which can be explained by a pragmatic phenomenon, i.e. scalar inference (si). to address whether this ambiguity patterns with that of scalars or numbers, our experiment 1 explored the behaviours of ‘a couple’ and scalars with two types of probe questions in inference tasks, and experiment 2 continued this theme by testing the naturalness rating for ‘a couple’ and scalars in an ‘x so not y’ construction. the results of our experiments indicate two natures of ‘a couple’: a non-monotonic /cardinal (approximately two) and proportional (a small proportion of). keywords. paucal quantifier; scalar inference 1. introduction. often hearers go beyond the literal meaning of what speakers utter and make inferences to enrich the message during language comprehension. scalar inferences (sis) are a widely discussed example of this kind of phenomenon: (1) a. some of the players scored. b. all of the players scored. c. not all of the players scored. regardless of whether thought of as grammatical or pragmatic, it is assumed that scalar inference is an operation which can augment the meaning of an utterance beyond what can be derived from its literal meaning. the operation involves the exclusion of a licenced alternative proposition (horn 1972, fox & katzir 2011). in the case of (1a), we assume that the literal meaning of the noun phrase ‘some of the players’ is a monotone increasing, existential quantifier function. then the sentence in (1b) can be an alternative and its exclusion leads to the implication in (1c). thus, the sentence in (2a) can optionally be understood as implying (2c). going beyond the classic example, much recent research has looked at a wider range of cases which can be given a similar analysis (van tiel et al. 2016, van tiel & schaeken 2017). when it comes to noun phrases (nps) which contain numerals, we can detect a similar duality of readings. consider (2): (2) a. spain has scored two goals b. if spain has scored two goals, they have turned the tide of the match. c. according to the scoreboard, spain has scored two goals. the sentence in (2a), when embedded in different contexts in (2b) and (2c), can be understood in different ways. in (2b) we could gloss the understanding as a lower-bounding, ‘two or more goals’, while in (2c), the gloss would be an ‘exactly’ reading – ‘two goals and no more’. if we extend the standard account of ‘some’ to numerals, the numeral ‘two’, when understood to have a literal meaning in the ‘two or more’ sense, would account for (2b). then a suitable alternative would be a sentence with ‘three’ and the result of combining the literal meaning with the alternative’s exclusion is the ‘exactly’ reading, prominent in (2c). this ‘standard’ account of the numeral case has long been challenged due to the sense that nps with * erying qin, university college london (erying.qin.18@ucl.ac.uk) chao sun, peking university (chaosun@pku.edu.cn) & richard breheny, university college london (r.breheny@ucl.ac.uk). proceedings of elm 3: 299-307, 2025 c©2025 erying qin, chao sun, richard breheny published by the lsa with permission of the author(s) under a cc by license. 299 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ numerals behave differently to those with ‘some’ and other expressions which are thought to be open to strengthening through scalar inferencing (horn 1992, geurts 2006, breheny 2008). moreover, a growing body of experimental research , e.g. marty et al. (2014), sun & breheny (2022), points to differences in outcomes for tasks when ‘some’ and numerals like ‘two’ are compared and these results are in line with the view that nps with numerals do not receive an ‘exactly’ reading as a result of strengthening by scalar inferencing. in this paper, we ask how the paucal quantifier involving ‘a couple’ behaves in terms of the likelihood of si. for example, ‘a couple’ can mean more or less the same thing as ‘two’, but it is also often used for a broader range of cardinal values than just two, depending on certain properties the objects in question are perceived to have. consider (3): (3) a. she scored (only) a couple of goals b. (only) a couple of fans in the crowd cheered. in (3a), the suggestion is that the quantity of goals scored was just two. in (3b), the quantity could be more, with the main implication being that the number is low relative to some standard. regarding the latter reading marty & nevins (under review) report studies showing that, in relatively neutral contexts, participants are prepared to judge quantities far higher than two as counting as ‘a couple’, as long the proportion is low relative to the whole. marty & nevins report similar results when the explicit operator ‘only’ is used, suggesting that a willingness to accept ‘a couple’ with larger numbers is not a result of a monotonic, ‘at least a couple’ meaning. one way to capture these data, supported by marty & nevins, is to assume that the quantity denoted by ‘a couple’ is only fixed relative to a context (to be a relatively small number). on this analysis, nps containing ‘a couple’, like ‘some’ are best analysed as existential, upward monotone quantifiers. however, one still needs to account for the fact that many participants judge sentences like in (3a,b) as false when the intersection of restrictor and scope has a cardinality larger than the (contextually determined) small number – as shown in marty & nevins. this suggests an accessible upper bounded meaning (‘a couple but not many’). marty & nevins assume this upper-bounded meaning arises through scalar inferencing with the alternative being ‘many’. an alternative would be to align ‘a couple’ with numerals and argue that the two kinds of interpretation of ‘a couple’ nps arise via different mechanisms. in our investigation of paucal quantifiers, we capitalize on various experimental paradigms to address the question whether ‘a couple’ patterns with ‘some’ or numerals. in the following part of the paper, we present two experiments: experiment 1 is based on sun & breheny’s (2022) inference tasks (section 2); experiment 2 is based on sun et al.’s (2018) ‘so’ task, which measures the naturalness of an si-enriched meaning under negation (section 3). 2. experiment 1. sun & breheny (2022) tested how the quantifier scale , the modal scale , and the numerical scale are interpreted, and established that genuine scalars ‘some’ and ‘possible’ are sensitive to a manipulation that can change the contextual relevance of alternatives (‘all’ and ‘certain’), whilst ‘exactly’ readings of numbers are not. sun & breheny (2022) investigated numerals, ‘possible’ and ‘some’ in inference tasks with two types of probe questions. one type, referred to as ‘not alt’ probe, was intrinsically a standard inference task where the probe question asked participants whether they could infer the negation of a scalar alternative (e.g. not all), according to a speaker character’s statement containing a scalar expression (e.g. some), and target response corresponding to inferring the si was a ‘yes’ response. the other type of probe question, called ‘could alt’ probe, asked participants whether, for instance, ‘all’ might not be excluded for the same statement, and target response was a ‘no’ response. note that participants could also give a ‘no’ response when they were uncertain about the speaker character’s intended meaning, irrespective of the probe type. in light of figure 1 (sun & breheny, 2022; p. 9, figure 4), the interpretations of proceedings of elm 3: 299-307, 2025 erying qin, chao sun, richard breheny: does ‘a couple’ pattern with scalars or numbers insights from inference and ‘so’ tasks. 300 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ‘some’ and ‘possible’ were affected by the manipulation of probes, because there were more target responses for ‘not alt’ than ‘could alt’ probes. figure 1: percentages of target responses for each scale and probe type (sun & breheny, 2022, p.9) this suggested that probe questions had an effect on making the si contextually relevant, so participants were more certain about inferring the si as part of the intended meaning, which led to more target responses for the ‘not alt’ probe. responses to numbers in their study showed a reversed pattern, indicating that making the contextual question more salient had no effect on target rates. sun & breheny explain the reverse pattern of results for numerals as resulting from the fact that the perceived ambiguity for numerals simply leads to an across the board increase in back-off ‘no’ responses due to uncertainty. this leads to more target responses for ‘could alt’ and fewer for ‘not alt’ – the attested pattern. our experiment 1 mirrors this study so as to see which effect manipulating contexts has on interpreting ‘a couple’. 2.1. participants. 60 native speakers of english participated in an online experiment run on gorilla experiment builder (anwyl-irvine et al. 2018). participants were recruited from prolific academic and compensated £0.8. all of them were naïve to the purpose of the experiment. participants were provided with an electronic version of informed consent before taking part, and this experiment was approved by the ucl research ethics committee. 2.2. materials and procedure. this experiment was a 2×3×3 inference task (probe type × condition × scale), and we manipulated condition and scale within subjects but probe type between subjects (we will elaborate these three factors below). to avoid the possibility that interpretations of ‘a couple’ are influenced by characteristics of numerals, we used the experimental items with to substitute those with numerals in sun & breheny’s (2022) original study (see figure 2). figure 2: examples of ‘not alt’ (left) and ‘could alt’ (right) probes the current experiment tested three scales: , and . the reason why we chose ‘many’ as the alternative is that when accounting for the two readings of ‘a couple’ in terms of si-based approach, marty & nevins (under review) proceedings of elm 3: 299-307, 2025 erying qin, chao sun, richard breheny: does ‘a couple’ pattern with scalars or numbers insights from inference and ‘so’ tasks. 301 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ regarded ‘many’ as its scalar alternative. for each scale, we constructed target, ‘true’ control and ‘false’ control conditions, and there were 6 items for target condition and 12 items for control condition (6 ‘true’ control items and 6 ‘false’ control items). in short, control items had the same structure as target items, except for the conclusions, the responses of which in the control conditions were either clearly ‘yes’ or clearly ‘no’. as probe type was a between-subject factor, each participant was randomly assigned to one of the two probe types and saw 54 items, including 6 target items, 6 ‘true’ control items and 6 ‘false’ control items per scale. two lists, the list of ‘not alt’ probe and that of ‘could alt’ probe, were created. each item only appeared once in each list, and the order of items was randomised for each participant in each trial. the inference tasks started with instructions and four practice trials. 2.3. results and discussion. participants were removed if their accuracy on control items was below 70%. three participants in the ‘not alt’ group and nine participants in the ‘could alt’ group were removed, and the overall mean accuracy of the control items reached 93% (‘true’ control condition: 89%, ‘false’ control condition: 97%). putting aside the control condition, we coded the ‘yes’ response to the ‘not alt’ probe and the ‘no’ response to the ‘could alt’ probe as target response. figure 3 shows the percentages of target responses for each scale and probe type. to analyse these target responses, we constructed a mixed effects logistic regression model predicting responses (target vs. nontarget) on the basis of probe type (not alt vs. could alt), scale type (‘some’ vs. ‘possible’ vs. ‘a couple’), and their interaction, including random intercepts for participants. random slopes were dropped due to non-convergence or singularity. the mixed-effect analyses, as well as all of the following analyses that will be reported, were conducted in r (r core team 2022) using the ‘lme4’ package (bates et al. 2015). degrees of freedom and corresponding p-values were estimated using the satterthwaite’s method, as implemented in the ‘lmertest’ package (kuznetsova et al. 2017). scale was dummy-coded, with ‘some’ as the reference level, and probe was deviation coded. model comparisons were conducted to test the significance of fixed effects with more than two levels, using likelihood ratio tests. significant interactions were followed up by conducting analyses on subsets of data defined by the levels of relevant factors. figure 3: percentages of target responses for each scale and probe type. error bars represent standard errors proceedings of elm 3: 299-307, 2025 erying qin, chao sun, richard breheny: does ‘a couple’ pattern with scalars or numbers insights from inference and ‘so’ tasks. 302 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the interaction between scale and probe type was significant (p < .001). on the target responses, there was a main effect of scale (p = .03) and a main effect of probe type (p < .001), indicating that the differences in the probability of target responses among three scales were greater for ‘not alt’ than ‘could alt’ probes. further analyses revealed that for ‘a couple’, the probability of target responses was higher for the ‘not alt’ probe compared to the ‘could alt’ probe (p < .01). the same effect was only marginally present for ‘possible’ (p = .07), but there were no statistically significant differences between the ‘not alt’ and the ‘could alt’ probes for ‘some’ (p = .23). our results were consistent with the pattern in sun & breheny between ‘not alt’ and ‘could alt’ probes for ‘possible’. in terms of ‘a couple’, the significantly different probability of target responses between the ‘not alt’ probe and the ‘could alt’ probe suggested that paucal quantifiers, such as ‘a couple’, behave like a genuine scalar expression, not like numbers. if we look at the data more closely, however, for ‘a couple’, the rate of target responses to ‘not alt’ probe was significantly lower than that for ‘possible’ (p = .04), which was similar to that found by sun & breheny when numerals were compared to scalars in ‘not alt’ trials . for the ‘could alt’ probe, the rate of target responses was lower for ‘a couple’ than for the other two scalars (some: p < .001; possible: p < .001) and there was no statistically significant difference between ‘possible’ and ‘some’ (p = 0.4). sun & breheny argue that the lower rate in ‘not alt’ trials for numerals than scalars is indicative of the fact that at least some participants recognized the ambiguity of the target trial sentence and became more non-committal. a similar drop off in rates for ‘a couple’ compared to the other scalars might indicate a kind of ambiguity between a non-monotonic, ‘exactly/approximately two’ interpretation and a proportional, ‘a small number of’ interpretation. the latter interpretation, like ‘some’ and ‘a few’, may be open to scalar inference, while the former, like numerals, not so. 3. experiment 2: naturalness rating for scalar expressions under negation. one distinguishing feature of numerals compared to many other widely discussed scalar expressions lies in their behaviour in linguistic contexts which tend to block scalar inference, such as in the scope of negation (horn 1992, breheny 2008). to illustrate, while (4a) is readily accepted in a case where she ate more than two cookies (say, three), (4b,c) are not readily acceptable where the stronger term is true; i.e. where it is certain she ate cookies in the case of (4b) or where she ate both a cookie and a cake in the case of (4c): (4) a. she did not eat two cookies. b. it is not possible she ate cookies. c. she didn’t eat a cookie or a cake. sun et al. (2018) report a study from which they extract a measure of felicity of scalar inference strengthening under negation. as per (4a) above, numeral expressions are quite felicitously understood in a non-monotonic sense in the scope of negation, while ‘possible’ and ‘or’ are not. sun et al. devised a ‘s so not w’ probe in order to collect felicity judgements for this. the idea behind the probe is that the sentence would be infelicitous unless there is local strengthening under negation. the ‘s’ term is simply an alternative that unilaterally entails the putative literal meaning of a monotone scalar expression. for example, for ‘possible’ we can use ‘certain’, as in (5b) below: (5) a. she ate three cookies, so not two. b. it is certain that she ate cookies, so not possible. c. the weather is hot, so not warm. assuming that scalar strengthening is blocked, or strongly disfavoured, under negation, the resulting ‘s so not w’ sentence should appear incoherent. but, to the extent that the nonmonotonic meaning is permissible under negation, the sentence should strike participants as proceedings of elm 3: 299-307, 2025 erying qin, chao sun, richard breheny: does ‘a couple’ pattern with scalars or numbers insights from inference and ‘so’ tasks. 303 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ less infelicitous (see breheny 2008 for more discussion). intuitively, this seems to be the case with (5a) above. sun et al. employed 42 different scalar terms, including ‘possible’, ‘warm’ and all others taken from van tiel at al. (2016). sun et al. took graded responses to ‘s so not w’ sentences as a measure of ‘liability to strengthen under negation’ and found that this measure can account for previously unexplained variance in rates of target responses in their replication of van tiel et al’s inference task study. sun et al.’s stimuli did not include numerals or ‘a couple’, and the highest rated scalar expressions on their ‘s, so not w’ task were only at the mid-range of a seven-point likert scale (see figure 4 below). in our second experiment, we wanted to bring numerals into the stimulus set and to explore more the idea that ‘a couple’ might have two kinds of interpretation, one that is more like ‘some’ and other scalar expressions, and one that is more like numerals. in order to do this, we re-considered what might be the alternative expression to use for ‘a couple’ in the stimuli. recall that in experiment 1, we used ‘many’ as the alternative for ‘a couple’, following on from marty & nevins. intuitively, the felicity of ‘many’ as an alternative for ‘a couple’ relies on the quantifier being understood in its proportional sense (see partee 1989). as discussed in relation to experiment 1 above, one way to account for the results would be to suppose that there is a second sense for ‘a couple’ which is more like ‘exactly/approximately two’. when thinking about the ‘s, so not w’ probe for ‘a couple’ we assumed that if the ‘s’ term were a partitive form involving ‘many’ then that would prime the proportional, scalar meaning, while if a non-partitive numerical np played the role of ‘s’, that would better prime any small-number approximative sense. these ideas were implemented in the design below. if we are right about participants having these two ways to understand ‘a couple’, we expected to see different outcomes using the different expressions in the ‘s’ role. specifically, when ‘s’ is a numeral, then felicity of ‘a couple’ should be more like that of numerals, compared to when ‘many’ is used. 3.1. participants. 103 native speakers participated for £1.5 compensation. recruitment and screening were identical to experiment 1. 3.2. materials and procedure. we used 48 scalars including 43 of them investigated in sun et al. ’s (2018) study along with numerals, ‘a couple/ number’, ‘a couple/many’ and some other scalars to construct experimental sentences for experiment 2. the experimental sentences were of the form ‘x so not y,’ where x and y were chosen according to the principle outlined above. figure 4 is an example item for ‘some’: figure 4: example displays for ‘some’ in the case of ‘a couple’ the two kinds of item are illustrated in figure 5 below: figure 5: examples of partitive (left) and non-partitive (right) groups we employed a between-subject design involving a partitive group and a non-partitive group. all the scalars in the two groups were the same, except for ‘a couple/number’ in the proceedings of elm 3: 299-307, 2025 erying qin, chao sun, richard breheny: does ‘a couple’ pattern with scalars or numbers insights from inference and ‘so’ tasks. 304 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ non-partitive group, whilst ‘a couple/many’ in the partitive group. each participant was randomly assigned to one of the two groups and judged 47 experimental sentences. participants were asked to indicate how natural these constructions are on a 1 (very unnatural) 7 (very natural) likert scale. in addition to experimental sentences, participants also had a chance to use the extremes of the likert scale for felicitous fillers (e.g. the window is open so not closed.) and infelicitous fillers (e.g. the train arrived so it never departed.). 3.3. results and discussion. one participant in the non-partitive group was excluded because the mean ratings for the infelicitous/filler items were above 5. again, the data analysis was performed using r (r core team 2022). the overall results for the two groups here compared with scales in sun et al. (2018) are shown in figure 6. when comparing our outcomes to those in the original study, we observed no significant difference in the mean ranks (wilcoxon signed-rank test: p = 0.089), thus the ‘so’-task results broadly align with those in sun et al. (2018). we note that overall, ratings for the 43 items that were common between sun et al’s previous study and this study were lower here. we assume this is due to the fact that participants tend to fix their range of ratings in relation to items that they have already seen and in this new experiment, the two new items involving numerals and ‘a couple’ were generally much more felicitous, making the other items seem less felicitous as a result. figure 6: mean naturalness ratings for ‘so’ task. note that ratings shown for all scales except for ‘a couple’ is the average across the two groups. proceedings of elm 3: 299-307, 2025 erying qin, chao sun, richard breheny: does ‘a couple’ pattern with scalars or numbers insights from inference and ‘so’ tasks. 305 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ turning now to the scalar expressions of interest, figure 7 shows the ratings from each group for stimuli with numerals, ‘a couple’ and three quantifier expressions which were highly ranked in both groups (and in sun et al’s previous study). we can see that, as expected, ‘so’task probes with numerals had a high felicity rating – higher than ‘a few’, ‘few’ and ‘some’. the question of interest for us is the comparative felicity of the probes for numerals and ‘a couple’ between the two groups. to compare the behaviour of ‘a couple’ and numerals in the two conditions, the mann whitney u test was conducted. we find a significant difference, when comparing ‘a couple/many’ to number (p = .05). however, there was no significant difference between ‘a couple/high number’ and number (p = 0.7). thus, the different expressions playing the ‘s’ role in the probe for ‘a couple’ had the expected effect. participants who saw a number like ‘20’, their judgement about the felicity of the probe was not different to the probe for numeral. when participants saw a sentence like ‘many..., so not a couple’ they found the sentence significantly less felicitous than the numeral items. figure 7: mean naturalness ratings for ‘so’ task in the partitive (left) and the non-partitive (right) group 4. summary. experiment 1 provides mixed evidence that ‘a couple’ patterns with genuine scalar items such as ‘some’; however, lower rate of target response to ‘a couple’ in ‘not alt’ condition compared to other scalars was similar to that of numerals in previous research. based on these results, we speculated that participants may have more than one way to interpret noun phrases with ‘a couple’. one way is more like those with ‘some’ and other existential monotone increasing quantifiers. one way is more like numerals, which have widely been viewed as not behaving in this way. turning now to experiment 2, we wanted to exploit the known felicity of the nonmonotonic interpretation of numerals in the scope of negation as a means to explore the possibility of two readings for ‘a couple’. using the ‘so’-task developed in sun et al. (2018), we devised two probe stimuli for ‘a couple’ each of which we expected to emphasise a different one of the two proposed ways that these noun phrases may be understood. the results show that when primed by a numeral alternative, ‘a couple’ behaved essentially in the same way as numerals in sun et al.’s ‘so’-task. when primed with ‘many’, participants found the probes less felicitous. overall, our results are suggestive that ‘a couple’ may semantically have two aspects. we acknowledge that this interpretation of the results is somewhat indirect and other interpretations are possible. we also leave open here more detailed analysis of how to formally analyse each of these two aspects of ‘a couple’ and to explain how they may be generated. references anwyl-irvine, al, jessica massonnié, adam flitton, natasha kirkham & jo k. evershed. 2020. gorilla in our midst: an online behavioral experiment builder. behavior research methods 52. 388–407. https://doi.org/10.3758/s13428-019-01237-x. proceedings of elm 3: 299-307, 2025 erying qin, chao sun, richard breheny: does ‘a couple’ pattern with scalars or numbers insights from inference and ‘so’ tasks. 306 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixedeffects models using lme4. journal of statistical software 67(1). 1–48. https://doi.org/10.18637/jss.v067.i01. breheny, richard. 2008. a new look at the semantics and pragmatics of numerically quantified noun phrases. journal of semantics 25(2). 93–139. https://doi.org/10.1093/jos/ffm016. fox, danny & roni katzir. 2011. on the characterization of alternatives. natural language semantics 19 (2011). 87–107. https://doi.org/10.1007/s11050-010-9065-3. geurts, bart. 2006. take ‘five’: the meaning and use of a number word. non-definiteness and plurality 95. 311–329. https://doi.org/10.1075/la.95.16geu. horn, laurence r. 1972. on the semantic properties of logical operators in english. ucla dissertation. horn, laurence r. 1992. the said and the unsaid. semantics and linguistic theory. 163–192. kuznetsova, alexandra, per b. brockhoff & rune h.b. christensen. 2017. lmertest package: tests in linear mixed effects models. journal of statistical software 82(13). 1–26. https://doi.org/10.18637/jss.v082.i13. marty, paul, and andrew nevins. under review. expressions of paucity: where is the upperbound?. ms ucl partee, barbara. 1989. many quantifiers. proceedings of the 5th eastern states conference on linguistics 5. 383–402. r core team. 2022. r: a language and environment for statistical computing. sun, chao & richard breheny. 2022. the role of alternatives in the interpretation of scalars and numbers: insights from the inference task. semantics and pragmatics 15(8). 1–15. https://doi.org/10.3765/sp.15.8. sun, chao, ye tian & richard breheny. 2018. a link between local enrichment and scalar diversity. frontiers in psychology 9. https://doi.org/10.3389/fpsyg.2018.02092. van tiel, bob, emiel van miltenburg, natalia zevakhina & bart geurts. 2016. scalar diversity. journal of semantics 33(1). 137–175. https://doi.org/10.1093/jos/ffu017. van tiel, bob & walter schaeken. 2017. processing conversational implicatures: alternatives and counterfactual reasoning. cognitive science 41. 1119–1154. https://doi.org/10.1111/cogs.12362. proceedings of elm 3: 299-307, 2025 erying qin, chao sun, richard breheny: does ‘a couple’ pattern with scalars or numbers insights from inference and ‘so’ tasks. 307 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ parenthesized modifiers in english and korean: what they (may) mean carolyn jane anderson & yoolim kim* abstract. although the semantics of some classes of parentheticals are well studied, such as appositives, there is relatively little work on parentheticals marked with parentheses. lewen & anderson (2022) analyze the semantics of a certain parenthesized construction they refer to as a restricted parenthesized parenthetical, and propose that the parentheses invoke and negate an alternative to the parenthesized content. this paper presents experimental evidence about the interpretation of parenthesized modifiers in korean and english, manipulating syntactic position and modifier properties (scalar/non-scalar, categorical/continuous). in both languages, our results confirm lewen & anderson (2022)’s proposal that some alternative is negated; however, the impact of the modifier properties we explore is different in english and korean. our findings corroborate the richness of the (often neglected) semantico-pragmatic space of parenthesized content. keywords. parentheses; parentheticals; alternatives; korean 1. introduction. existing work on parentheticals has focused on their use as appositives, speakeroriented adverbials, and expressives (mccawley 1982, ziv 1985, potts 2002, dehé & kavalova 2007). there is little existing work on parentheticals that are marked with parentheses (nunberg 1990) or on parentheticals outside of indo-european (kim 2012). this paper presents experimental evidence about the interpretation of one kind of parenthesized parenthetical in both american english and korean. we focus on a parenthesized construction discussed in lewen & anderson (2022), which they refer to as a restricted parenthesized parenthetical (rpp). this construction gives rise to an implication that its non-parenthesized counterpart does not, as shown in (1). (1) a. sam studies linguistics for (intellectual) profit. # and actual profit. b. sam studies linguistics for intellectual profit. and actual profit. (lewen & anderson 2022) lewen & anderson (2022) highlight key differences between this construction and better-studied classes of parentheticals like appositives. they propose an analysis in which the parentheses act as a focus-sensitive operator, invoking and negating a set of alternatives to the parenthesized content. in this paper, we test their hypothesis that rpps invoke and negate alternatives experimentally in a dialogue interpretation task. we explore the effect of two semantic properties of the parenthesized modifier and present a cross-linguistic comparison between american english and a language with different syntactic restrictions on parentheticals: korean. a key difference between korean and english is that in korean, the parenthesized parenthetical can come on either side of the modified noun, as in (2), while in english rpps, it must be on *thank you to hangyeol park and jin ryu for help in translating and running the korean experiment. authors: carolyn jane anderson, wellesley college (carolyn.anderson@wellesley.edu) & yoolim kim, wellesley college (yoolim.kim@wellesley.edu). proceedings of elm 3: 1-18, 2025 c©2025 carolyn jane anderson and yoolim kim published by the lsa with permission of the author(s) under a cc by license. 1 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the left (nunberg 1990). it is not known whether the meaning contribution of rpps in korean is similar to english, or whether the two syntactic positions correspond to differences in meaning. (2) a. sam-neun sam-top (cicek) (intellectual) iik-ul gain-acc wuyhay for enehak-ul linguistics-acc kongpu-ha-p-ni-ta study-do-ah-ind-decl sam studies linguistics for (intellectual) profit. b. sam-neun sam-top iik-ul gain-acc (cicek) (intellectual) wuyhay for enehak-ul linguistics-acc kongpu-ha-p-ni-ta study-do-ah-ind-decl sam studies linguistics for (intellectual) profit. we test this experimentally in dialogue interpretation task, where participants judge whether alternatives to the parenthesized content are excluded in the context. we manipulate two key semantic properties of the modifiers: whether they are scalar or non-scalar; and whether they are continuous or categorical. we also explore the syntactic position of the parenthesized component in korean. our cross-linguistic comparison reveals that although the korean and american english constructions appear similar syntactically, their semantics are not identical: while both languages are consistent with lewen & anderson (2022)’s analysis of rpps as negating alternatives, the effect of modifier properties differs between the two languages. these findings highlight the need for more work exploring the fine-grained meaning contribution of parentheses cross-linguistically. 2. restrictive parenthesized parentheticals. parentheticals comprise a broad class of linguistic phenomena, including appositives, speaker-oriented adverbials, expressives, and more. what unites them is a sense that they contribute “extra” information or commentary beyond the basic meaning contribution of the sentence. a key property of parentheticals is independence: they can be removed or omitted without affecting their host sentence. much previous work has focused on understanding the extent to which parentheticals are syntactically and semantically independent (mccawley 1982, potts 2002, 2005, blakemore 2006, dehé & kavalova 2007, de vries 2007, dehé 2009, blakemore 2009, mcinnerney 2020). mccawley (1982) argues that parentheticals are not always syntactically independent of their hosts: in (3), for instance, the object of the verb sells in the parenthetical bill knows a man who sells is the host sentence object. (3) mary buys, and bill knows a man who sells, pictures of elvis presley. even in these cases, though, the host sentence is still independent, since the entire parenthetical could be deleted without affecting the grammaticality or interpretation of its host. in one of the few works that engages with parenthesized parentheticals, nunberg (1990) asserts that the host sentences of parenthesized parentheticals are always independent, writing that “the content of a parenthetical must be entirely irrelevant to the syntactic or semantic well-formedness of the surrounding text” (nunberg 1990; p. 106). however, lewen & anderson (2022) show that the host sentences of one kind of parenthesized construction, which they call the restrictive parenthesized parenthetical (rpp), are not independent. they provide corpus examples of the construction, such as (4) below, where removing the parenthetical results in ungrammaticality. (4) such a set would preserve the print and (some) of the tools used to create it. (davies 2008; taken from lewen & anderson (2022)) they also argue that the meaning of the parenthesized component and its host are more closely proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 2 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ intertwined than traditional parentheticals, which, as nunberg writes, leave behind “a coherent and complete communication” when removed (nunberg 1990; p. 106). for instance, removing the parenthesized component in (5) makes the host sentence infelicitious in the given context. (5) context: calvin has eaten some, but not all, of the cake. a. calvin didn’t eat (all of) the cake. b. # calvin didn’t eat the cake. (lewen & anderson 2022) thus, though rpps, like better-studied classes of parentheticals, seem to add “extra” information, they do not have the same degree of independence as other kinds of parentheticals. 2.1. parentheses as focus operators. lewen & anderson (2022) propose an analysis of the semantics of rpps centered around the infelicity of the continuation shown in (1-a) and (6-a). (6) a. sally drinks (herbal) tea before bed. # or black tea. b. sally drinks herbal tea before bed. or black tea. they argue that the parentheses in an rpp work like a focus operator in that they invoke alternatives to parenthesized content. in their proposal, the infelicity arises because the parentheses negate an alternative to their contents. in (1-a), the most contextually salient alternative is financial gain. this negated alternative conflicts with the continuation, leading to infelicity for the rpp, but not its non-parenthesized equivalent. in the lewen & anderson (2022) analysis, rpps assert their non-parenthesized equivalent and presuppose the negation of a contextually salient alternative to the parenthesized component. their proposed semantics are shown in (7). (7) semantics of the rpp construction (lewen & anderson 2022): [[α(β)]]c = a. asserts: αβ b. presupposes: ∃alt′ ⊆ altc(β).∀δ∈alt′ .(¬αδ) ∧ (δ >c β) where altc takes a constituent γ and returns a set of relevant alternatives to γ and >c is an alternative strength criterion in c. example (8) shows how their analysis treats (6-a). (8) meaning contribution of (6-a) according to (lewen & anderson 2022) a. asserts: sally drinks herbal tea before bed. b. presupposes: {¬sally drinks white tea before bed,¬sally drinks oolong tea before bed, ¬sally drinks green tea before bed,¬sally drinks black tea before bed} the rpp in (6-a) contributes the assertion that sally drinks herbal tea, as well as the presupposition that sally does not drink some other kind of tea. given the context, a plausible set of alternatives to be excluded are alternatives that involve more strongly caffeinated tea. in this paper, we test lewen & anderson (2022)’s proposal that rpps negate alternatives, and explore a question that lewen & anderson (2022) leave open: how the alternative(s) are selected. 2.2. alternatives in rpps. lewen & anderson (2022) propose that the parentheses in rpps proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 3 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ negate alternatives to their content. however, they leave open the set of alternatives involved, both in terms of its size and its selection, providing examples in which all alternatives seem to be negated (9), as well as examples where some alternatives do not seem to be excluded, such as (10). (9) often sally drinks (herbal) tea before bed. # or black tea. # or green tea. (lewen & anderson 2022) (10) context: for health reasons, arjun tries not to eat red meat, and he also dislikes turkey. a. # arjun avoids only red meat. b. arjun avoids (red) meat. (lewen & anderson 2022) in (10), there is a contrast between the rpp and only, which is known to negate all alternatives to its associate (horn 1969, rooth 1985). one possibility is that it is only stronger alternatives that are negated. however, it is not clear how to define strength: lewen & anderson (2022) explore and reject both logical strength and comparative likelihood. our experiment seeks to shed light on alternative exclusion by manipulating two semantic properties of the parenthesized modifiers: whether they are scalar, and whether their category boundaries are categorical or continuous. if alternative strength is relevant, this may be easiest to observe with scalar modifiers, which have a natural ordering. on the other hand, if only a single alternative is negated, this may be easiest to observe with categorical modifiers, where the boundaries between alternatives are most clear. 2.3. cross-linguistic comparison. our work seeks both to test lewen & anderson (2022)’s analysis of american english rpps, and to investigate the cross-linguistic stability of the construction’s meaning contribution. we compare the interpretation of rpps in american english with korean, a language with little existing research on parentheticals (kim 2012). korean offers an interesting comparison because it provides a syntactic alternation: the parenthesized modifier can appear either to the left or right of the noun (11), even though it can only appear on the left in the non-parenthesized version (12).1 (11) a. sally-neun sally-top cak-ijeney sleep-before cha-lul tea-acc (tay-chwu) (jujube) ma-sin-ta drink-ind-decl sally drinks tea (herbal) before bed. b. sally-neun sally-top ca-ki-jen-ey sleep-before (tay-chwu) (jujube) cha-lul tea-acc ma-sin-ta drink-ind-decl sally drinks (herbal) tea before bed. (12) a. sally-neun sally-top ca-ki-jen-ey sleep-before tay-chwu jujube cha-lul tea-acc ma-sin-ta drink-ind-decl sally drinks herbal tea before bed. b. *sally-neun sally-top ca-ki-jen-ey sleep-before cha-lul tea-acc tay-chwu jujube ma-sin-ta drink-ind-decl * sally drinks herbal tea before bed. in english, by contrast, rpps can only appear to the left. this was first observed by nunberg 1although the translation notes ‘herbal’, we use a more culturally-relevant tea, ’jujube tea’, which is often had before sleep. this substitution is meant to provide a comparable context with a less caffeinated tea. proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 4 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ a: are you still doing a lot of volunteer work for the pet shelter? b: i don’t do as much as i used to, but i still write their (weekly) newsletter. question: which kinds of newsletters do you think b doesn’t help to write? ( ) daily ( ) monthly ( ) other: figure 1: example item from the scalar categorical condition. (1990), who distinguishes two classes of parenthesized parentheticals: constituents of lexical phrases, like lewen & anderson (2022)’s rpps, that appear to the left of their heads, and parentheticals belonging to what he calls the text-grammar, which must come after their heads: (13) a. *they include (as they put it) “free gifts” with every purchase. b. they include “free gifts” (as they put it) with every purchase. (nunberg 1990) thus, in english, syntactic position demarcates two classes of parenthesized parentheticals, with parentheticals that are tightly integrated into their lexical phrases appearing to the left. in korean, both syntactic positions are possible for rpps; however, it is not known whether their interpretation differs based on their position. for korean, therefore, we can explore the meaning impact of the syntactic manipulation as well as of the different classes of modifiers. 3. experimental design. to determine whether rpps negate alternatives to their parenthesized content, we ran two forced-choice dialogue interpretation tasks. experiment a explores the interpretation of rpps in english and experiment b explores their interpretation in korean. 3.1. procedure. both experiments used a dialogue interpretation paradigm with two interlocutors, a and b. the first interlocutor poses a question, and the second replies with a sentence containing an rpp. the participant was then asked to select from a list of alternatives to the parenthesized content all alternatives that are not compatible with the given context. for instance, in the item shown in figure 1, b’s utterance contains the rpp (weekly) newsletter. the answer options contain two alternatives to weekly, daily and monthly. participants were asked to judge which of these options, if any, are ruled out by b’s utterance. participants could select one or more options. they could also provide a different response using the other option. 3.2. conditions. we manipulated two properties of modifiers: scalarity some modifiers naturally fall onto a scale. for instance, warm represents the midrange on a temperature scale, between hot and cold. some scales may be highly salient in certain contexts. for instance, in the context of drinking tea right before bed, the caffeine level of each kind of tea becomes salient. we call these modifiers scalar. on the other hand, other modifiers do not easily form an ordered scale. for instance, the set of flavors for a particular candy may fall into discrete categories, but these categories do not occur in any order. similarly, cardinal directions in english are conventionally listed with south and north “bordering” east and west, but the listing proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 5 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ can begin at any position; these modifiers form a cycle rather than a scale. we call these kinds of modifiers non-scalar. we classify modifiers into two groups based on whether there is a salient scale available for ranking them. for scalar modifiers, we parenthesize a middle element from the scale in order to test whether alternative strength matters in the interpretation of rpps. for non-scalar modifiers, we simply select two alternatives. continuity another modifier property that may be relevant is the nature of the category boundaries. we group modifiers based on whether they demarcate clear and exclusive category boundaries. for instance, levels of school (elementary, middle, high) are discrete categories with clear boundaries: a student is not typically in both kinds of schools at once. we call modifiers like these categorical. on the other hand, the boundaries between times of day may be more negotiable. for instance, between morning and afternoon, we could create additional distinctions like late morning, early afternoon, or brunch time. we call modifiers like these continuous. size size modifiers are common in rpps, but challenging to categorize by our criteria. although size is clearly scalar, in some contexts, size boundaries are fuzzy, while in others, they are conventionalized. for instance, t-shirts and to-go beverage cups come in discrete size categories, but pets do not, leading to some degree of ambiguity when apartment listings allow small dogs only. for this reason, we include size modifiers as a separate condition. we use medium as the parenthesized modifier for all items in the size condition, and select small and large as the alternative options. syntactic position in korean, we tested an additional manipulation of position: the parenthetical appeared either to the right or the left of the modified noun. 3.3. items. we crossed scalarity and continuity to create four modifier conditions: scalar categorical, scalar continuous, non-scalar categorical, and non-scalar continuous. for each modifier condition, we constructed ten dialogue sets, each with a different modifier within the category. we also constructed six dialogues with size rpps. appendix b gives an example of each condition. for korean, we presented half of the items in each syntactic position. thus, english participants judged ten items in each main condition and six size items, for a total of 46 items; korean participants judged five main and three size items for each syntactic position, for a total of 46 items. we also included three training items and three filler items. the training included an item where none of the options were appropriate, to model the use of the other option (appendix a). 3.4. participants. data was collected from 32 native korean speakers and 32 monolingual english speakers. we collected information about the languages spoken in each participant’s childhood households and the language of instruction in the school they attended most recently. we used this information to exclude participants with significant exposure to a language other than the target language. 3.5. free response data coding. participants were allowed to select an other option and write in a response. these responses were coded into the following categories: • none: none of the options are excluded; all of them are possible in the context. • all: all of the options listed are excluded; any other option would be excluded. • high: something stronger than what is parenthesized is excluded. • low: something weaker than what is parenthesized is excluded. proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 6 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ • justify: the answer simply justifies the selected options. • specificother: mentions a specific alternative that is excluded (other than a or b). • focusother: suggests an alternative to an element other than the target modifier. • undecidable: the context does not provide enough information to decided. • parenthesized: the parenthesized item is excluded. • irrelevant: the response is unclear or unrelated to core question. the same coding categories were used for korean and english responses. examples of responses in each category can be found in appendix c. we consider parenthesized and focusother as error cases, where participants misunderstood the task or context. 3.6. statistical analysis. we fit mixed-effects logistic regression models to the response data from each experiment. we fit models for each language for each of the three main response patterns: selecting both alternatives (both), selecting a single alternative (single), or other. we also fit a model to the pooled korean and english data to explore between-language differences. 4. results. 4.1. experiment a: english. experiment a explores the interpretation of rpps in english. according to the analysis put forward by lewen & anderson (2022), rpps invoke and negate an alternative to their parenthesized content. if this is correct, we expect participants to choose at least one of the alternatives (or to suggest their own). if participants do not select any alternatives or they write that all are possible, this would indicate that rpps do not necessarily negate an alternative. experiment a explores five classes of modifiers. we manipulate whether the modifier is scalar or not and whether it is continuous or categorical. we include size as a separate category, because depending on context, it can be either continuous or categorical. 4.1.1. main conditions. overall, we find that most participants select at least one alternative as excluded given the dialogue (88%). this is consistent with lewen & anderson (2022)’s hypothesis that rpps negate an alternative to their parenthesized content. 4.1.2. scalar and non-scalar conditions. we observe different response patterns between modifier conditions. figure 2 shows the response selection patterns for the four main modifier conditions. there is a clear difference between scalar and non-scalar conditions in the distribution of other responses. in non-scalar conditions, the majority of participants select both alternative options (49.3% both; 18.9% high; 15.5% low), indicating that multiple alternatives are negated. in the scalar conditions, by contrast, the most common response type is to select only one of the alternatives (36.9% high; 33.6% low; 19.7% both). there was a significant negative effect of scalarity in the english both mixed-effects model and a significant positive effect of scalarity in the english single mixed-effects model (appendix d). we do not see as strong of an effect from manipulating the nature of the categories, though category type seems to interact with scalarity. in the scalar continuous condition, there are fewer both responses than in the scalar categorical condition, while in the non-scalar continuous condition, there are even more both responses than in the non-scalar categorical condition. the english both and english single mixed-effects models find no significant primary affect of categorical, but confirm that its interaction with scalarity is statistically significant (appendix d). proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 7 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 2: english option selections by modifier type. error bars indicate 95% cis. figure 3: by-item responses for english scalar conditions one surprising finding is that it is not always the stronger alternative that is selected, when only one option is chosen: in the scalar categorical case, there are roughly even rates for both options. one possibility is that while participants are sensitive to the presence of a scale, they differ its ordering. perhaps in low cases, participants are using the expected scale, but reversing its direction. when we examine the response patterns in the scalar conditions by item (figure 3), we find that by item, most participants agree on which alternative is excluded. however, there is variability between items in whether it is the stronger or weaker alternative. this supports the hypothesis that a stronger alternative that is negated: if the alternative selection were arbitrary, we would expect items with roughly equal numbers of strong and weak selections. our paradigm is not designed to differentiate between excluding a single stronger alternative or excluding all stronger alternatives, since the scalar answer options consist of two alternatives on opposite ends of the scale. however, the free response data provides some tentative evidence of participants excluding all stronger alternatives. ten of the responses indicate that any stronger alternative would be excluded and four indicate that any weaker alternative would be excluded.2 4.1.3. size. the size condition was included separately from the other modifiers because it can be hard to categorize as either continuous or categorical. figure 3 shows the response patterns for the size condition. as expected, size patterns most similarly to the scalar conditions. however, there is a somewhat larger proportion of both responses, suggesting that it is closer to the scalar 2for example, for an item with the scalar categorical rpp (chapter) books, one participant selected the stronger alternative, textbooks, and additionally wrote in, “books above the level of chapter books.” proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 8 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ category english korean none 114 51 all 21 12 high 10 5 low 4 4 justify 6 0 specificother 17 9 focusother 1 10 undecidable 37 86 parenthesized 5 9 irrelevant 12 3 table 1: free responses by language and coded category. figure 4: english results, including coded free response answers categorical condition than the scalar continuous. 4.1.4. free responses. the experimental design allowed participants to write in an option. participants used this in a variety of ways: to provide justification for their choice (coded justify); to specify additional excluded alternatives (specificother); to indicate that all alternatives would be excluded (all), or that no alternatives would be excluded (none), or that there was not enough information to decide (undecidable). some responses were unclear or off-topic (irrelevant). overall, english participants used the other option 227 times (for 1426 items). the most common response category was none, followed by undecidable (table 1). figure 4 shows the response patterns when the coded free responses are included. we combine the high category with responses where only the stronger alternative was checked (only a); the low category with only b; and the all category with responses where both options were checked (both). the english other mixed-effects model finds a significant effect of categorical on other responses and a significant interaction with scalar, suggesting that participants had more trouble deciding for these categories (appendix d). however, the rate of none and undecidable responses is low in all conditions, suggesting that most participants were sensitive to rpp alternatives. 4.2. experiment b: korean. experiment b explores the same modifier conditions as the english experiment, but with an additional manipulation of syntactic position. we manipulate proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 9 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 5: option selections by language, modifier type, and position. error bars indicate 95% cis. whether the parenthesized modifier appears to the right or to the left of its noun. figure 2 shows the korean response selection patterns for the four main modifier conditions by syntactic position, compared with english. as in english, we find that most participants select at least one alternative (82%). however, the rate of other selection is higher than in english. the responses pattern very similarly for both right and left placement of the parenthetical, indicating that syntactic position is not associated with differences in meaning; none of the mixedeffects models find a statistically reliable effect of position (appendix d). the effect of modifier property is not as clear. in general, the response patterns are similar to english, but weaker, with large numbers of participants using each response strategy in every condition. the korean both mixed-effects model finds a significant positive effect of scalarity and a marginal negative interaction with categorical (appendix d). similarly, the korean single mixedeffects model finds a significant negative effect of scalarity and a significant positive interaction with categorical, echoing the english trends. however, although there are more both responses in the non-scalar conditions, it is not the most common response in any category: korean participants most frequently select only one of the options across categories. the cross-language mixed-effects models confirm that this is a statistically reliable difference between the languages (appendix d). 4.2.1. free responses. korean participants used the free response option more frequently than english participants. they were also more than twice as likely to use this option to express that they did not have enough information to select any of the options (undecidable). this was the most common free response category for korean speakers (table 1). like english speakers, the none response category was also common. when we combine the free responses with the option selection patterns, we again see few differences by syntactic position. for both positions, we see more undecidable responses in the non-scalar categories than the scalar categories, as well as a more mixed pattern of options selections, with substantial numbers of participants selecting both options, only a, and only b. although participant behavior is more mixed in the korean experiment, the rate of participants proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 10 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 6: korean results by syntactic position, including coded free response answers figure 7: difference between korean only a/high and only b/low counts by condition. responding with a none free response is consistently low across conditions, suggesting that korean participants are sensitive to rpps invoking some alternatives. an alternative explanation would be that korean participants are more likely to write that they cannot decide rather than to write that no options are excluded. if this is the case, then the higher numbers of undecidable responses may indicate that the alternative negation effect of rpps is less strong in korean. 4.2.2. alternative strength. as in english, we observe no clear preference for selecting the stronger alternative. figure 7 shows the difference between stronger and weaker alternative selection counts. the preference for the stronger or weaker option is consistent by item, but does not pattern with modifier category, suggesting that it may be an artifact of the modifiers used in each item. this is consistent with the hypothesis that scale reversal is involved in cases where participants exclude only the weaker alternative. 5. discussion. experiment a and experiment b explored the interpretation of rpps in english and korean. overall, the results of both experiments confirmed lewen & anderson (2022)’s hypothesis that rpps exclude alternatives to their parenthesized content. in both experiments, most participants selected at least one alternative to exclude. the experiments also tested the impact of two modifier properties, scalarity and continuity. in proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 11 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ english, there is a strong effect of scalarity: in non-scalar conditions, most participants exclude both alternatives, while in scalar conditions, participants typically select only one alternative to exclude. surprisingly, participants did not always select the strongest alternative in scalar conditions; we hypothesize, however, that this may be due to by-item cases of scale reversal. the differences between continuous and categorical modifiers were less strong; instead, continuity seems to interact with scalarity to intensify its effect, with non-scalar continuous modifiers receiving the most both responses and scalar continuous modifiers receiving the fewest both responses. when we turn to the korean results, we observe a similar but weaker effect of scalarity. in general, the korean response patterns are more mixed across all modifier conditions. although more participants excluded both options in non-scalar conditions compared to scalar conditions, across conditions, most participants excluded only one option. moreover, korean participants were more likely to say that the context did not provide enough information to decide against any of the alternatives. this suggests that the interpretation of rpps in korean is different from english, or at least less consistent: there seems to be more variability in how korean rpps are interpreted. overall, our results provide evidence in support of lewen & anderson (2022)’s theoretical account of rpps for english, while illustrating the need for more cross-linguistic analysis of parenthetical constructions. the results of experiment b provide a first look at the interpretation of parentheticals in korean and suggest that more work is necessary to understand the meaning contribution of rpps in this language. 6. conclusion. we present experimental evidence about the interpretation of a particular kind of parenthesized parenthetical, the restrictive parenthesized parenthetical. we test lewen & anderson (2022)’s hypothesis that rpps invoke and negate at least one alternative to their parenthesized content, and explore the nature of the alternative(s) targeted. in experiment a, we explore the interpretation of rpps in english and investigate two key properties of the parenthesized modifiers. in experiment b, we provide the first experimental evidence of how rpps are interpreted in korean, and explore the impact of syntactic position along with modifier properties. our results support lewen & anderson (2022)’s hypothesis that rpps involve the negation of at least one alternative to their parenthesized content. we also answer an open question posed in lewen & anderson (2022) about the nature of the alternative set: we find that english rpps with non-scalar modifiers negate all of their contextually relevant alternatives, while those with scalar modifiers negate alternates at one end of their scale. finally, our results illustrate the need for more cross-linguistic work on the semantics of parentheticals: although our korean data are consistent with the alternative-negation account of rpps, the effect of modifier type is not clear-cut. we hope that future work will shed more light on cross-linguistic differences in the fine-grained semantics of parentheticals. references blakemore, diane. 2006. divisions of labour: the analysis of parentheticals. lingua 116(10). 1670–1687. blakemore, diane. 2009. on the relevance of parentheticals. in actes d’idp, . davies, mark. 2008. the corpus of contemporary american english (coca): 600 million words, 1990-present. available online at www.english-corpora.org/coca. proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 12 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ de vries, mark. 2007. invisible constituents? parentheses as b-merged adverbial phrases. in nicholas dehé & yordanka kavalova (eds.), parentheticals, 203–234. amsterdam: john benjamins. dehé, nicholas. 2009. clausal parentheticals, intonational phrasing, and prosodic theory. journal of linguistics 45. dehé, nicholas & yordanka kavalova. 2007. parentheticals: an introduction. in nicholas dehé & yordanka kavalova (eds.), parentheticals, 1–24. benjamins. horn, laurence. 1969. a presuppositional analysis of only and even. chicago linguistic society 5. 98–107. kim, mija. 2012. syntactic types of as-parentheticals in korean. in stefan müller (ed.), proceedings of the 19th international conference on head-driven phrase structure grammar, chungnam national university daejeon, 216–231. stanford, ca: csli publications. 10.21248/hpsg.2012.13. lewen, carina bolaños & carolyn jane anderson. 2022. (some) parentheses are focus-sensitive operators. in proceedings of sinn und bedeutung, vol. 26, 165–186. mccawley, james d. 1982. parentheticals and discontinuous constituent structure. linguistic inquiry 13(1). 91–106. mcinnerney, andrew. 2020. parentheticals associate with their hosts pragmatically, not syntactically: evidence from as-parentheticals. in mariam asatryan, yixiao song & ayana whitmal (eds.), nels, vol. 50, 177–187. amherst, ma: glsa. nunberg, geoffrey. 1990. the linguistics of punctuation. csli publications. potts, christopher. 2002. the syntax and semantics of as-parentheticals. natural language and linguistic theory 20. 623–689. potts, christopher. 2005. the logic of conventional implicatures oxford studies in theoretical linguistics. oxford university press. rooth, mats. 1985. association with focus: university of massachusetts, amherst dissertation. ziv, yael. 1985. parentheticals and functional grammar. in machtelt bolkestein, caspar de groot & j. lachlan mackenzie (eds.), syntax and pragmatics in functional, 181–199. dordrecht: de gruyter mouton. appendix a training instructions in this study, you’ll see a short dialog between two speakers, a and b, like the one below: a: what do you think i should wear to the party tonight? b: i’m planning to wear a polo. you will be asked a multiple choice about the dialog. you may have to draw some inferences about what a and b mean in order to answer the question. for instance, here is an example question about the previous dialog: proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 13 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ question: based on b’s information, which of the following shirts do you think a might consider wearing? ( ) a polo ( ) a button-down shirt ( ) other: you can check more than one option, if you think more than one of the shirts listed would be appropriate for a to wear. for instance, you might infer that the party is not casual enough to wear a tee shirt, but a might be willing to wear either a polo or something slightly dressier. or, you might think that a wants to avoid being too formal, so both the tee shirt and polo are possible. we’re interested in your interpretation of the dialogs. so you should pick however many options you think might work. if none of the options seem good to you, you can write ”none” in the other category or suggest a different option. a: are you thinking about getting a pet? b: i’m allergic to everything. question: which of the following pets do you think b might adopt? ( ) a cat ( ) a dog ( ) other: here’s one last training item for you to practice on: a: do you think mark will be ok if i make kimchi fried rice for dinner? i know he’s picky. b: he eats (seafood) fried rice. question: which of the following kinds of fried rice do you think mark will eat? ( ) kimchi ( ) egg ( ) other: b example stimuli from each modifier condition scalar categorical a: how is lulu’s reading progressing? b: she doesn’t read (chapter) books yet. question: which books do you think lulu can’t read? proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 14 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ( ) textbooks ( ) picture books ( ) other: scalar continuous a: what would you like to have with the steaks? b: they would go really nicely with some (moderately) roasted carrots. question: which cooking styles does b think wouldn’t go with the steaks? ( ) lightly ( ) very dark ( ) other: non-scalar categorical a: do you think liza will like these candies? b: i know that she likes (strawberry) hi-chew. question: which flavors do you think liza might not like? ( ) mango ( ) apple ( ) other: non-scalar continuous a: i hear you’re planning a trip to europe! where are you hoping to go? b: i’m passionate about wine, so i’m planning a route through france and (southern) germany. question: which regions of germany do you think b won’t visit? ( ) western ( ) northern ( ) other: size a: i’m so excited for our trip to denmark! how much luggage are you planning to bring? b: just my backpack and one (medium) suitcase. question: which sizes of suitcase do you think b would not bring? ( ) small ( ) large ( ) other: c free response coding an example of a free response entry for each category from the english responses is shown below. • none: “liking strawberry doesn’t imply disliking any other flavor.” • all: “any jacket other than denim.” • high: “any cardigan made with heavy material.” • low: “distances less than 5 miles.” proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 15 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ • justify: “all dogs might be allowed, because medium is kind of vague.” • specificother: “suit.” • focusother: “dogs off the leash.” • undecidable: “not enough information to know.” • irrelevant: “i’m not familiar with blanched or stir fried asparagus.” d mixed-effects model results we fit mixed-effects logistic regression models to three response variables: both selection (1 if both options are selected or there was a free response in the both or high category, and 0 otherwise); single selection (1 if exactly 1 of the options is selected or there was a free response in the high or low category, and 0 otherwise); and other selection (1 if neither of the given options were selected and the free response was not in the both, all, high, or low categories, and 0 otherwise). each model includes fixed effects of modifier scalarity (1 for scalar conditions and 0 for nonscalar conditions) and continuity (1 for categorical conditions and 0 for continuous conditions). the korean models additionally include syntactic position (a = left, b = right). interaction terms were included fof the fixed effects as well. models also included random effects for participants and items, with maximal random effects structures. fixed effects β̂ z p intercept -0.8 (+/0.2) -3.6 0.0003 scalar 2.3 (+/0.2) 11.4 <0.0001 categorical -0.006 (+/0.2) -0.03 0.98 scalar*categorical -0.71 (+/0.3) -2.4 0.01 table 2: mixed-effects english single model fixed effects β̂ z p intercept 0.17 (+/0.2) 0.7 0.48 scalar -2.8 (+/0.3) -10.6 <0.0001 categorical -0.42 (+/0.3) -1.6 0.11 scalar*categorical 1.6 (+/0.3) 4.7 <0.0001 table 3: mixed-effects english both model proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 16 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ fixed effects β̂ z p intercept -2.4 (+/0.3) -7.2 <0.0001 scalar 0.3 (+/0.3) 1.0 0.32 categorical 0.59 (+/0.3) 2.3 0.02 scalar*categorical -1.1 (+/0.4) -2.8 0.005 table 4: mixed-effects english none model fixed effects β̂ z p intercept 0.72 (+/0.4) 1.8 0.08 scalar 0.97 (+/0.3) 2.9 0.004 categorical 0.18 (+/0.3) 0.6 0.52 position b -0.09 (+/0.3) -0.3 0.76 scalar*categorical -0.70 (+/0.4) -1.7 0.095 scalar*position b 0.66 (+/0.5) 1.4 0.17 categorical*position b -0.32 (+/0.4) -0.8 0.43 scalar*categorical*position b -0.05 (+/0.6) -0.08 0.9 table 5: mixed-effects korean single model fixed effects β̂ z p intercept -1.57 (+/0.4) -3.6 0.0002 scalar -1.18 (+/0.4) -2.8 0.005 categorical -0.55 (+/0.3) -1.7 0.09 position b 0.24 (+/0.3) 0.76 0.45 scalar*categorical 1.34 (+/0.5) 2.5 0.01 scalar*position b -0.63 (+/0.6) -1.1 0.29 categorical*position b 0.0008 (+/0.4) 0.002 1.0 scalar*categorical*position b 0.35 (+/0.8) 0.46 0.65 table 6: mixed-effects korean both model fixed effects β̂ z p intercept -2.9 (+/0.5) -5.9 <0.0001 scalar -0.52 (+/0.4) -1.2 0.25 categorical 0.87 (+/0.4) 2.3 0.02 position b -0.16 (+/0.4) -0.38 0.71 scalar*categorical -0.02 (+/0.5) -0.03 0.98 scalar*position b -0.04 (+/0.7) -0.06 0.95 categorical*position b 0.46 (+/0.5) 0.92 0.36 scalar*categorical*position b -0.58 (+/0.8) -0.69 0.49 table 7: mixed-effects korean none model proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 17 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ fixed effects β̂ z p intercept -0.85 (+/0.3) -2.9 0.004 scalar 2.4 (+/0.2) 11.2 <0.0001 categorical 0.07 (+/0.2) 0.35 0.73 korean 1.39 (+/0.4) 3.6 0.0004 scalar*categorical -0.80 (+/0.3) -2.8 0.005 scalar*korean -1.09 (+/0.3) -3.4 0.0006 categorical*korean 0.01 (+/0.3) 0.03 0.97 scalar*categorical*korean 0.17 (+/0.4) 0.4 0.69 table 8: mixed-effects cross-language single model fixed effects β̂ z p intercept 0.18 (+/0.3) 0.62 0.54 scalar -2.9 (+/0.3) -10.5 <0.0001 categorical -0.39 (+/0.2) -1.9 0.06 korean -1.43 (+/0.4) -3.7 0.0002 scalar*categorical 1.7 (+/0.3) 5.2 <0.0001 scalar*korean 1.5 (+/0.4) 3.7 0.0002 categorical*korean -0.31 (+/0.3) -1.0 0.32 scalar*categorical*korean -0.37 (+/0.5) -0.8 0.45 table 9: mixed-effects cross-language both model fixed effects β̂ z p intercept -2.4 (+/0.4) -6.8 <0.0001 scalar 0.08 (+/0.3) 0.3 0.77 categorical 0.57 (+/0.3) 2.2 0.03 korean -0.08 (+/0.5) -0.2 0.87 scalar*categorical -0.78 (+/0.4) -2.1 0.03 scalar*korean -0.35 (+/0.4) -0.9 0.39 categorical*korean 0.13 (+/0.4) 0.4 0.70 scalar*categorical*korean 0.37 (+/0.5) 0.7 0.49 table 10: mixed-effects cross-language none model proceedings of elm 3: 1-18, 2025 carolyn jane anderson and yoolim kim: parenthesized modifiers in english and korean: what they (may) mean. 18 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ relating scalar inference and alternative activation: a view from the rise-fall-rise tune in american english thomas sostarics, eszter ronai, and jennifer cole* abstract. the rise-fall-rise (rfr) tune in american english has received numerous theoretical accounts to describe its meaning contribution, with a consistent theme being the relationship between rfr and “higher alternatives.” however, autosegmentalmetrical theory predicts three rfr-shaped tunes which differ in the rising pitch accent used (h*, l+h*, l*+h), raising the question of whether different rfr-shaped tunes in fact behave differently. we investigate this question under the lens of scalar inference (si). we find that rfr-shaped tunes with different pitch accents behave similarly in offline interpretation, increasing the rate of si calculation relative to falling tunes. in online processing using cross-modal priming with lexical decision, we find an asymmetry in the processing profile of two rfr-shaped tunes: h*l-h% leads to additional facilitation of the higher alternative, while l*+hl-h% leads to less facilitation. we describe these results in relation to differences in pitch range and discuss how they relate to ongoing debates about rfr. keywords. intonation, prosody, scalar inference, priming, rise-fall-rise 1. introduction. the rise-fall-rise (rfr) tune in american english has received ample theoretical attention in semantics and pragmatics over the past forty years. the meaning contribution of rfr has been variably described as conveying uncertainty with regard to a higher value along a scale (ward & hirschberg 1985, hirschberg & ward 1992); conveying the existence of disputable higher alternatives (constant 2012); highlighting the salience of a higher alternative (göbel 2019, göbel & wagner 2023); indirectly or partially addressing a question under discussion (qud) (wagner et al. 2013) or relating to a (hierarchically higher) contrastive topic or secondary qud (büring 2003, westera 2019). recent empirical work on rfr (de marneffe & tonhauser 2019, buccola & goodhue 2023, ronai & göbel to appear) has sought evidence for or against these accounts using experimental methods, finding variation in how adequately different theoretical proposals account for their results. while accounts of rfr will often describe a singular rfr, typically1 referencing the (tobi annotated) l*+hl-h% contour described by ward & hirschberg (1985), autosegmental-metrical (am) phonological theory predicts not one but three putatively distinct rfr-shaped tunes which differ in the choice of rising pitch accent: h*, l+h*, or l*+h (pierrehumbert 1980). are the three rfr-shaped tunes interpreted similarly, indicating a broad class of rfr-shaped tunes, or are they interpreted differently from one another? we investigate potential contrasts between the rfr-shaped tunes under the lens of scalar inference (si) in both *we thank kate sandberg for help with recording, michael tabatowski for feedback on the written stimuli, and chun chan for experiment implementation support. we also thank the prosd lab and experimental meaning group at northwestern and the audiences at the voices in contexts workshop and experiments in linguistic meaning for feedback on this work. authors, all at northwestern university: thomas sostarics (tsostarics@u.northwestern.edu) & eszter ronai (ronai@northwestern.edu), & jennifer cole (jennifer.cole1@northwestern.edu). 1 proceedings of elm 3: 383-394, 2025 c©2025 thomas sostarics, eszter ronai, and jennifer cole published by the lsa with permission of the author(s) under a cc by license. 383 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ offline interpretation and online processing. the conceptual link between rfr and si is as follows. when describing the various pragmatic accounts for rfr, what stands out across accounts (rhetorically, though not necessarily implementationally) is a persistent invocation of some notion of “higher alternatives.” si is famously another domain within semantics/pragmatics where higher alternatives play a role. in the (neo-)gricean tradition, si is taken to arise via reasoning about what a speaker could have said but did not (grice 1975) with particular attention to pairs of lexical items that form a lexical scale (i.a., horn 1972). such scales are described in terms of relative informativity and can be formalized using a relation of asymmetric entailment (horn 1972) such that for a scale , an utterance containing y entails one containing x but not the other way around; hence, y is the informationally stronger member of the pair and is often referred to as the stronger or higher scalemate while x is the weaker or lower scalemate. in (1), some is the weaker scalemate of the scale, leading the listener to reason about the speaker’s use of some instead of all, ultimately arriving at the si-enriched interpretation of jane ate some but not all of the cookies from the literal meaning of the sentence.2 (1) jane ate some of the cookies. jane ate some, and possibly all, of the cookies. literal jane ate all of the cookies. alternative ⇝ jane ate some but not all of the cookies. si-enriched in the context of si computation, some accounts of rfr make competing predictions regarding whether si should be more or less likely when a sentence like (1) is uttered with rfr. specifically, if the use of rfr conveys uncertainty about whether a higher alternative y is true or not, then si—the negation of y —would be incompatible with such uncertainty. for instance, one cannot be uncertain whether jane ate all the cookies while simultaneously concluding that jane did not eat all the cookies. recent experimental work on rfr has also relied on si as its empirical testing ground, and has found that the use of rfr increases the likelihood of si calculation (de marneffe & tonhauser 2019, ronai & göbel to appear, though cf. buccola & goodhue 2023).3 these results thus appear prima facie as evidence against uncertainty accounts of rfr, but it remains an open question whether a different pattern of results, potentially in line with uncertainty accounts, might be obtained with a different rfr-shaped tune. 1there is some variation in how researchers describe rfr. for example, büring (2003; 537) describes an (l+)h*lh% tune for contrastive topic marking, which constant (2012; 431) argues is distinct from l*+hl-h%. westera (2019; 326) notes a potential difference between h*l-h% and l*+hl-h%, but nonetheless elects to group the two together when comparing büring and constant’s (among others’) accounts of rfr to one another. similarly, wagner et al. (2013; 130) annotates rfr with an l+h* accent, attributing this to hirschberg & ward (1992) who specifically differentiate l*+h from l+h* in the context of rfr (ibid 242, see also ward & hirschberg 1985; 748). regardless, the historical development of this literature has sought to identify the meaning contribution of “rfr,” but it is not clear when, or whether, such accounts are expected to hold for putatively different rfr-shaped tunes. 2the likelihood that such si-enriched interpretations arises varies across lexical scales, termed the scalar diversity phenomenon (van tiel et al. 2016). in work on scalar diversity, a primary goal is to understand the systematic factors (semantic, pragmatic, or otherwise) that contribute to this variation in rates of si computation (gotzner et al. 2018, ronai & xiang 2021, 2024, a.o.). the present work does not seek to explain scalar diversity, but rather uses it as a testing ground for competing predictions from existing accounts of rfr. 3while de marneffe & tonhauser (2019; 12) interpret their empirical results as showing that rfr “strengthens 2 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 384 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ in addition to probing the offline interpretation of different rfr-shaped tunes, one might also wonder whether there is a psycholinguistic processing correlate sensitive to the relationship between rfr and higher alternatives (and again, whether this varies for different rfr-shaped tunes). we test this question using cross-modal priming with lexical decision, which has previously been used to identify the activation of focus alternatives attributable to different intonational features (braun & tagliapietra 2010, husband & ferreira 2016, gotzner et al. 2016). specifically, contrastive focus alternatives (rooth 1992) show sustained activation later in processing while words that cannot serve as contrastive focus alternatives, yet are nonetheless semantically related to the focus-accented element in a sentence, become deactivated later in processing. in our cross-modal lexical decision experiment, we extend prior text-only priming studies investigating scalar alternatives (ronai & xiang 2023, lacina & gotzner 2024) to not only investigate whether higher scalar alternatives are activated similarly to focus alternatives but, crucially, whether rfr has a modulating effect on this activation. in summary, discussion of rfr often brings with it discussion of higher alternatives. higher alternatives also play a key role in si. our goal is to use si as a testing ground to identify potential differences in interpretation between three putatively contrastive rfr-shaped tunes. we present experimental results from an inference judgment task, used to assess offline interpretation, and a cross-modal lexical decision task, used to assess online processing of rfr. 2. materials. we use both cross-modal priming with lexical decision and the inference judgment task (henceforth just “inference task”) using the same set of auditorily-presented materials. rfr cannot be used out of the blue, so we need to have a preceding context before an utterance with rfr. accordingly, we wrote indirect polar question-answer (q/a) pairs between two speakers like in (2), where bob’s response uses either the lower alternative cool or the higher alternative cold. (2) alice: did someone leave a window open in the office overnight? bob: the office feels cool / cold. these q/a pairs differ from those used in de marneffe & tonhauser (2019) because alice’s question does not use the higher alternative cold. this constraint on the materials is needed due to the lexical decision task: if we are interested in the activation status of cold following bob’s reply “the office feels cool.”, then having cold explicitly mentioned in the preceding utterance (alice’s question) will directly activate it and likely mask any potential priming effect from intonation (see gotzner et al. 2016 and yan & calhoun 2019 for effects of mentioned vs unmentioned alternatives in lexical decision). lastly, these contexts are written to neither bias towards nor against si calculation when bob’s answer contains the lower alternative. a literal interpretation of the office feels cool (and possibly cold) may be taken to be a positive response to the question (i.e., yes someone left a window open, explaining the chilliness) and an si-enriched interpretation of the office feels cool (but not cold) may be taken to be a negative response to the question (i.e., no, nobody left a window open), but si calculation itself is not required to arrive at a yes or no response (though, a the degree of belief in the scalar implicature,” ronai & göbel (to appear; 447) interpret their findings, which also show an increase in si computation, as rfr enhancing the salience of the higher alternative (göbel 2019, göbel & wagner 2023), where increased salience of the higher alternative has been shown (through other means) to make si more likely (ronai & xiang 2021). other accounts may be compatible with either an increase or decrease in si rates, e.g., constant (2012), westera (2019). 3 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 385 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ relevance implicature is indeed needed). in total, we wrote contexts for 74 adjectival scales reported in prior work (aparicio & ronai 2023, de marneffe & tonhauser 2019, van tiel et al. 2016) to use as our critical trial stimuli in both experiments. we also wrote an additional 122 filler trials, which comprise two types. first, 61 are q/a pairs with non-word targets in the lexical decision task—we will refer to these as the non-word fillers. second, to avoid a task adaptation effect where participants might learn that the only time they need to respond yes is when they encounter a sentence-final adjective, we also include 61 q/a pairs adapted from husband & ferreira (2016), example shown in (3), such that the accented word in the sentence (bolded) is placed at the end of bob’s response analogously to the critical trials—we will refer to these as the hf16-adapted items. (3) original item: the museum thrilled the sculptor when they called about his work. alice: did the museum deliver any good news? bob: the museum thrilled the sculptor. our critical materials were normed in a preliminary text-only task, which included naturalness rating and inference judgment components. based on the rating results we removed 10 scales from our task that received a high proportion of low ratings, leaving a total of 64 critical items. the remaining critical items were recorded by the first author in six intonation conditions: three rfr-shaped tunes (using l-h% edge tones) and three falling tunes (l-l%) that each differ in pitch accent (h*, l+h*, and l*+h). these recordings were then modeled using generalized additive models to create consistent targets for resynthesis (via psola in praat boersma & weenink 2020) for each tune in order to reduce the amount of variation across sentences. the averages of each tune across all the resynthesized recordings are shown in figure 1 and are evenly distributed across all items.4 3. inference task. we recruited 84 participants from the online crowdsourcing platform prolific for our inference task. this task is similar to prior work on si calculation (van tiel et al. 2016, ronai & xiang 2021) where, on each trial, participants listened to a pre-recorded dialogue such as (2), where the answer includes the lower alternative from a lexical scale such as . they were then presented with a question like would you conclude that the office does not feel cold? and had to answer with “yes” (indexing si calculation) or “no”. we included an additional 72 fillers, 36 of which come from the non-word fillers and 36 of which come from the hf16adapted items. the items were distributed into 12 counterbalanced lists and then presented in a pseudorandom order that minimizes adjacent trials having the same intonational tune or item set. each participant thus saw 136 trials (64 critical plus 72 fillers) divided into four blocks. 3.1. results. the average empirical si rates for each intonation condition are shown in figure 2. we can observe that the rfr-shaped tunes as a group show higher average si rates than the falling tunes with potential graded distinctions between the three based on the pitch accent used. we analyze these results with bayesian logistic mixed effect regression models using brms (bürkner 2021) with weakly informative regularizing priors. we include a fixed effect of tune, the combination of pitch accent (h*, lh*, l*h) and edge-tone configuration (ll, lh) and random 4while not shown here, our materials show previously described patterns of pitch accent peak height and alignment across accent types (iskarous et al. 2024); an extended phonetic discussion must be left to a future paper. 4 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 386 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 40 60 80 100 120 140 160 word start peak end normalized time pi tc h (h z) h*ll lh*ll l*hll lh*lh l*hlh h*lh figure 1: resynthesized utterances for each tune, time-normalized to the location of the pitch accent peak. the averages of each tune are superimposed on top of the individual contours and labeled at the location of the final f0 value on the right-hand side. intercepts by item and participant and random slopes of tune by participant and item. what we’re interested in is the overall effect of edge-tone configuration (i.e., rfr-shaped versus falling) and the degree to which there are graded distinctions across the pitch accents, which we operationalize in terms of the differences of lh*−h* and l*h−h* within each broad tune class (rfr and falling). we use a manually-specified contrast matrix to encode these comparisons in a single model. our analysis code is available at https://osf.io/bc6a2/. from the statistical model we find a main effect of edge-tone configuration (β̂ = 0.35, ci = [0.18, 0.52]) such that the rfr-shaped tunes together yield higher si rates than the falling tunes. the distinctions between pitch accents are small, and unsurprisingly given the overlap between the pitch accents evident in figure 2, the 95% credible intervals for each pitch accent comparison contain 0. although, the bulk of the posterior distribution for the l*hlh−h*lh comparison (β̂ = 0.18, ci = [−0.08, 0.45]) is greater than 0, with a probability of direction of 90.7%. these results replicate prior work on rfr in the context of si (de marneffe & tonhauser 2019, ronai & göbel to appear) but provide a novel finding that there is a primary dichotomy between broad falling and rfr-shaped tune classes. moreover, this effect persists even when the higher alternative is not explicitly mentioned in the question (c.f. ronai & göbel to appear a.o.). while there appear to be numerical differences in si rates between pitch accents within these broad classes, there remains ample uncertainty as to the magnitude and systematicity of such distinctions. this caveat is perhaps not entirely surprising in light of work on variation in intonational form showing overlap in intonational categories in both production and perception (arvaniti 2019, cole et al. 2023, steffman et al. 2024). such overlap has also presented itself in psycholinguistic investigation, where despite much work drawing a convenient categorical distinction between “neutral” h* and “focus-marking” or “contrastive” l+h*, contrastive interpretations are nonetheless possible even when h* is used (watson et al. 2008). with regard to prior accounts of rfr, 5 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 387 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 0.27 0.30 0.33 0.36 h*ll lh*ll l*hll h*lh lh*lh l*hlh tune a vg s i r at e figure 2: average empirical si rates for each tune with ±2 standard errors. our results are compatible with theoretical proposals from constant (2012), göbel (2019), westera (2019) but not accounts that require the truth of the higher alternatives to remain open through conveying uncertainty (ward & hirschberg 1985) or leaving the alternative unresolved or only partially addressed (wagner et al. 2013). 4. cross-modal priming. based on prior work on rfr, our primary hypothesis is that rfr evokes higher alternatives. in priming terms, we predict that when rfr is used with an utterance containing a lower alternative like cool, we will see facilitation in lexical retrieval of the corresponding higher alternative cold due to rfr boosting its activation level. furthermore, since si is uni-directional, we also expect that when rfr is used with a higher alternative, we will not see facilitation of the lower alternative; for instance, if rfr is used with cold, we might expect it to evoke a higher alternative like freezing but not a lower alternative like cool. whether the lower alternative is specifically inhibited or simply not affected by rfr is left unspecified. based on the results of exp. 1, we expect that the three rfr-shaped tunes will behave similarly, predicting facilitation of lexical retrieval of the higher alternative for all three tunes, relative to falling tunes. this task presents 186 trials split into six blocks of 31 trials. of these, 64 trials are critical items which vary by intonational tune and by whether the higher or lower alternative serves as the visual target (with the other scalemate serving as the auditory prime). we will refer to the condition where the higher alternative is the target (i.e., hear cool then see cold) as the highertarget condition and the condition where the lower alternative is the target (hear cold then see cool) as the lowertarget condition. the fillers include 61 non-word fillers and 61 real-word fillers form the hf16-adapted items. the six intonational tunes evenly distributed across each set of filler items. each hf16-adapted item has one of three possible targets; using an auditory prime of sculptor as an example, the three targets are a contrastive semantic associate (i.e., a contrastive focus alternative) like painter, a noncontrastive associate like statue, and a semantically unrelated word like register. we recruited 104 undergraduate students from northwestern university, who participated for 6 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 388 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ course credit. 41 participants were omitted due to language background or issues with completing the task, leaving 63 participants for analysis. on each trial, participants listened to a recorded audio dialogue. the visual target appeared on the screen 750ms after the offset of the final word of the utterance. participants were instructed to use a button box to judge whether this visual target is a word or not a word of english (framed as yes, it is a word, or no, it is not a word). the experiment was implemented in psychopy (peirce 2007) and administered in a sound-attenuated booth with a 165hz monitor; participants gave their responses using a cedrus rb-740 buttonbox. 4.1. results. before we investigate the effect of intonation on reaction times (rts), we first look at how our critical items compare to the hf16-adapted items. one condition in the hf16adapted items is the contrastive condition, where the target word can serve as a (contrastive) focus alternative to the prime word. similarly, our adjectival scalemates, e.g., cool and cold, can also serve as focus alternatives to one another. however, our critical items additionally comprise a lexical scale and are hence related via asymmetric entailment. accordingly, we are interested in first evaluating whether lexical scalemates offer any processing advantage beyond their status as focus alternatives. by doing so, we also gain a baseline for the rts in each condition (i.e., what can be attributed to the relationship between the lexical items of the auditory prime and visual target) before seeing how intonation further modulates these. we analyze participant rts in each condition when correctly responding yes (total accuracy 98.6%) using a bayesian lognormal distributional model. our main predictor of interest is target condition5 controlled for effects of log word frequency (balota et al. 2007) of the target word, length of the target word, and experimental block (all treated as continuous and centered). we include random intercepts by participant and item and random slopes of condition by participant; to account for differences in rt dispersion we also include random intercepts for the sigma parameter of the model by participant and item. the model-predicted rts are shown in figure 3. from the statistical model, we find that rts for higher alternatives are not credibly different from lower alternatives (β̂ = 0, ci = [−0.03, 0.02]). moreover, the rts for these scalar alternatives are not credibly different from contrastive alternatives (β̂ = 0.01, ci = [−0.01, 0.03]). rts for noncontrastive associates are slower than contrastive and scalar alternatives (β̂ = 0.05, ci = [0.02, 0.07]) while semantically unrelated words are slower than all semantically related words (β̂ = 0.09, ci = [0.06, 0.11])—this result shows the same pattern as the focus intonation condition in husband & ferreira (2016; 227). overall, these results suggest that when the visual target can serve as a contrastive alternative to the auditory prime, there is not an additional processing advantage when the target and prime words are additionally related via asymmetric entailment. given that the two critical conditions show strong facilitation compared to the semantically unrelated filler condition, does intonation modulate the degree of facilitation within each condition? that is, does intonation contribute anything beyond the priming induced by the visual prime being a possible scalar alternative? based on the hypothesis that rfr invokes higher alternatives and our offline results from exp. 1 showing that the rfr-shaped tunes behave similarly to one another, we predict that the three rfr-shaped tunes will lead to additional facilitation in the highertarget condition (i.e., cool uttered with rfr leads to faster lexical decisions of cold). moreover, 5condition has five levels and is helmert coded to encode nested orthogonal comparisons, which are depicted in figure 3. estimates are on the loge rt scale. 7 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 389 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 400 450 500 550 600 condition pr ed ic te d r t lowertarget cool highertarget cold contrastive painter noncontrastive statue unrelated register * * figure 3: model-predicted rts with 95% posterior credible intervals. rts are marginalized over target word length and frequency and experimental block. credible differences are shown by a ∗. because rfr is taken to specifically invoke higher alternatives, we predict that rfr will not lead to additional facilitation for lower alternatives (i.e., cold uttered with rfr should not lead to faster lexical decisions of cool). in other words, there should be an asymmetry in the processing profile of rfr given the directional relationship between the auditory prime and the visual target. we focus on only the critical lowertarget and highertarget conditions, again using a bayesian lognormal distributional model with an added predictor of tune and its interaction with condition.6 figure 4 shows the posterior predicted differences, in terms of percent change (%∆), of each tune-condition combination to the condition means previously seen in figure 3. we can observe that, overall, intonation does not have any notable affect within the lowertarget condition, which is in line with the prediction that rfr does not modulate activation of the lower alternative. within the highertarget condition, the falling tunes are also not credibly different from the highertarget condition mean. we do find credible evidence of additional facilitation for h*lh% (β̂ = −0.019,%∆ = −1.86%, ci = [−0.035,−0.003]), but less facilitation for l*hl-h% (β̂ = 0.019,%∆ = 1.96%, ci = [0.003, 0.036]). lh*l-h%, which lies in-between h*l-h% and l*hl-h% in phonetic space, does not show credible evidence for a difference from the conditional mean. generally, that any of the rfr-shaped tunes (here, 2 of them) show an asymmetry such that a change in processing is observed in the highertarget condition but not the lowertarget condition is in line with predictions that rfr is associated with higher alternatives. yet, unlike in the si rate data in exp. 1, here we do not observe a pattern wherein all rfr-shaped tunes behave alike; rather, focusing on h*l-h% and l*hl-h%, we have two tunes within the same 6here, condition is treatment coded with highertarget as the reference level (coded 0) and lowertarget as the comparison level (coded 1). tune is a 6-level predictor that is sum coded (h*l-l% used as the reference level, coded -1, and comparison levels coded as +1). together, these contrast schemes yield fixed effect estimates that reflect deviations of each tune from the highertarget condition’s mean. 8 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 390 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ lowertarget highertarget h* lh* l-l% l*h h* lh* l-h% l*h h* lh* l-l% l*h h* lh* l-h% l*h -5.0% -2.5% 0.0% 2.5% 5.0% tune % d iff er en ce fro m c on di tio n m ea n -0.36 +0.27 -0.23 +0.26 -0.01 +0.07 +0.05 -0.86 +0.33 -1.86 +0.42 +1.96 figure 4: posterior distribution of rt percent changes with 50%, 89%, and 95% mean highest density intervals. tunes that are not credibly different from the condition means are shown in gray (means=black circles) while tunes that are credibly different are shown in red (light diamonds). broad class behaving differently. regardless, the results of this online task are directly in line with the common theme throughout the rfr literature that rfr is related to higher alternatives, thus providing psycholinguistic evidence for this claim. 5. general discussion. we investigated whether rfr-shaped tunes behave differently from one another as compared to falling tunes using the same three pitch accents. in our offline inference task, we found a primary distinction between rfr-shaped tunes and falls, with the former encouraging the computation of si-enriched interpretations. this replicates prior empirical findings on rfr in the context of si computation (de marneffe & tonhauser 2019, ronai & göbel to appear) with not only a greater variety of intonational tunes, but also with dialogue contexts that do not overtly mention the higher alternative. (note that buccola & goodhue (2023) find different results, which may be attributable to differences in the experimental paradigms —see buccola & goodhue (2023) and ronai & göbel (to appear) for further discussion.) with regard to formal pragmatic accounts of rfr, our results are in line with proposals from göbel (2019) and göbel & wagner (2023) where the use of rfr enhances the salience of the higher alternatives. these results are incompatible with accounts of rfr which predict a reduction in the likelihood of si computation such as ward & hirschberg (1985) and wagner et al. (2013) but may be compatible with other accounts such as those offered by constant (2012) and westera (2019), which do not unambiguously predict either an increase or decrease in si computation. in our online cross-modal priming with lexical decision task, we find an asymmetry such that two rfr-shaped tunes show different processing signatures when probing a higher scalemate like cold compared to a lower scalemate like cool. interestingly, while we hypothesized that increased likelihood of si would lead to increased facilitation of the higher alternative, we only find this for 9 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 391 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ one rfr-shaped tune, h*l-h%, which was not the rfr that yielded the highest si rate. moreover, the rfr that did yield the highest si rate instead showed less facilitation of the higher alternative.7 our two sets of results seem to paint different pictures: rfr-shaped tunes behave alike in offline interpretation but behave (counterintuitively) differently in online processing. we can reconcile these results by considering variation at the level of the holistic tune, rather than between-category variation at the level of pitch accents. our priming results show a difference in the amount of facilitation between the rfr-shaped tunes that use h* (≈2% faster) versus l*+h (≈2% slower), with l+h* lying somewhere in the middle— so we still need to contrast, minimally, h*l-h% with l*+hl-h%. we interpret this pattern similarly to hirschberg & ward (1992), who report that increased scaling in the pitch range of rfr is more likely to yield an incredulous interpretation. more broadly, it is known that increased scaling of pitch range and pitch accent excursions lead to higher conveyance of speaker affect, engagement, and arousal (ladd et al. 1985, gussenhoven 2004). while our materials are distinct in their tonal specification (e.g., h*l-h%), they also co-vary in terms of their pitch range; l*+h is known to be more prominent, with higher and later-aligned peak f0 targets, compared to l+h* and h* (iskarous et al. 2024). accordingly, we can recast our rfr-shaped tunes under a broad rfr class (based on our si rates, where the rfr-shaped tunes behaved similarly) with meaningful phonetic variation between a low-scaled rfr (our h*l-h%) and a high-scaled rfr (our l*+hl-h%). when interpreting a high-scaled rfr, additional competing inferences beyond si (such as incredulity or other particularized inferences) may become more likely. in our priming task, this allows for an interpretation whereby rfr evokes higher alternatives, leading to facilitation, but further increasing the scaling of rfr invites additional competing inferences which strains processing and may mask potential facilitation effects, resulting in less facilitation. in our inference task, the response options are constrained to specifically probe si computation, allowing for differences in si rates to arise. 6. conclusions. we presented results from offline interpretation and online processing experiments, finding evidence for within-category variation at the level of broad falling and rise-fallrise tune shapes. in offline interpretation, we found that rfr-shaped tunes increase the likelihood of si computation compared to falls in indirect question-answer dialogue contexts. in online processing, we found that scalar alternatives behave similarly to contrastive focus alternatives. moreover, we find an asymmetry when rfr is used with a lower alternative, leading to either additional facilitation when the rfr is low-scaled (h*) or less facilitation when the rfr is high-scaled (l*+h). neither effect is found when probing the lower alternative after rfr is used with a higher alternative, which rules out the interpretation that the facilitation effect we do find for higher alternatives can be reduced to an effect of semantic priming. altogether, our results therefore provide evidence of a psycholinguistic correlate to the recurring theme in the literature that rfr evokes higher alternatives. one limitation of our priming experiment is that it is not known whether or not a participant computed the si-enriched interpretation, precluding directly linking the presence of 7independently, lacina & gotzner (2024) report a similar counterintuitive pattern in a text-based priming paradigm, where participants made a lexical decision on the higher alternative following rapid serial visual presentation of a sentence containing a lower alternative. the authors found an inverse correlation between si rates and facilitation across lexical scales such that scales with higher si rates displayed less facilitation of their higher alternatives; in other words, the priming results showed the inverse of the typical scalar diversity cline of si rates. 10 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 392 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ activation for alternatives to si calculation. in ongoing work, we address this issue by combining cross-modal lexical decision and inference tasks in a dual-task setup. references aparicio, helena & eszter ronai. 2023. scalar implicature rates vary within and across adjectival scales. in semantics and linguistic theory, 110–130. 10.3765/t7t8pn98. arvaniti, amalia. 2019. crosslinguistic variation, phonetic variability, and the formation of categories in intonation. in proceedings of the international congress of the phonetic sciences, canberra, australia: australasian speech science and technology association inc. balota, david a, melvin j yap, keith a hutchison, michael j cortese, brett kessler, bjorn loftis, james h neely, douglas l nelson, greg b simpson & rebecca treiman. 2007. the english lexicon project. behavior research methods 39. 445–459. 10.3758/bf03193014. boersma, paul & david weenink. 2020. praat: doing phonetics by computer [computer program]. version 6.2,14. braun, bettina & lara tagliapietra. 2010. the role of contrastive intonation contours in the retrieval of contextual alternatives. language and cognitive processes 25(7-9). 1024–1043. buccola, brian & daniel goodhue. 2023. the effect of intonation on scalar and ignorance inferences. in the proceedings of the chicago linguistic society, vol. 59, 1–12. büring, daniel. 2003. on d-trees, beans, and b-accents. linguistics and philosophy 26(5). 511– 545. 10.1023/a:1025887707652. bürkner, paul-christian. 2021. bayesian item response modeling in r with brms and stan. journal of statistical software 100(5). 1–54. 10.18637/jss.v100.i05. r package version 2.21.6. cole, jennifer, jeremy steffman, stefanie shattuck-hufnagel & sam tilsen. 2023. hierarchical distinctions in the production and perception of nuclear tunes in american english. laboratory phonology 14(1). 10.16995/labphon.9437. constant, noah. 2012. english rise-fall-rise: a study in the semantics and pragmatics of intonation. linguistics and philosophy 35(5). 407–442. 10.1007/s10988-012-9121-1. göbel, alexander. 2019. additives pitching in: l*+h signals ordered focus alternatives. in semantics and linguistic theory, vol. 29, 279–299. 10.3765/salt.v29i0.4612. göbel, alexander & michael wagner. 2023. on a concessive reading of the rise-fall-rise contour: contextual and semantic factors. experiments in linguistic meaning 2. 83–94. gotzner, nicole, s solt & a benz. 2018. scalar diversity, negative strengthening, and adjectival semantics. frontiers in psychology 9(sep). 1–13. 10.3389/fpsyg.2018.01659. gotzner, nicole, isabell wartenburger & katharina spalek. 2016. the impact of focus particles on the recognition and rejection of contrastive alternatives. language and cognition 8(1). 59–95. grice, herbert p. 1975. logic and conversation. in speech acts, 41–58. brill. gussenhoven, carlos. 2004. the phonology of tone and intonation. cambridge university press. hirschberg, julia & gregory ward. 1992. the influence of pitch range, duration, amplitude and spectral features on the interpretation of the rise-fall-rise intonation contour in english. journal of phonetics 20(2). 241–251. 10.1016/s0095-4470(19)30625-4. horn, laurence robert. 1972. on the semantic properties of logical operators in english: university of california, los angeles dissertation. 11 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 393 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ husband, e matthew & fernanda ferreira. 2016. the role of selection in the comprehension of focus alternatives. language, cognition and neuroscience 31(2). 217–235. iskarous, khalil, jennifer cole & jeremy steffman. 2024. a minimal dynamical model of intonation: tone contrast, alignment, and scaling of american english pitch accents as emergent properties. journal of phonetics 104. 101309. 10.1016/j.wocn.2024.101309. lacina, radim & nicole gotzner. 2024. exploring scalar diversity through priming: a lexical decision study with adjectives. in proceedings of the annual meeting of the cognitive science society, vol. 46, https://escholarship.org/uc/item/4mr322mh. ladd, d robert, kim ea silverman, frank tolkmitt, günther bergmann & klaus r scherer. 1985. evidence for the independent function of intonation contour type, voice quality, and f 0 range in signaling speaker affect. the journal of the acoustical society of america 78(2). 435–444. de marneffe, marie-catherine & judith tonhauser. 2019. inferring meaning from indirect answers to polar questions: the contribution of the rise-fall-rise contour. in questions in discourse, 132–163. brill. 10.1163/9789004378322 006. peirce, jonathan w. 2007. psychopy—psychophysics software in python. journal of neuroscience methods 162(1-2). 8–13. 10.1016/j.jneumeth.2006.11.017. pierrehumbert, j. 1980. the phonology and phonetics of english intonation: massachusetts institute of technology dissertation. ronai, eszter & alexander göbel. to appear. watch your tune! on the role of intonation for scalar diversity. glossa psycholinguistics . ronai, eszter & ming xiang. 2021. exploring the connection between question under discussion and scalar diversity. proceedings of the linguistic society of america 6(1). 649. ronai, eszter & ming xiang. 2023. tracking the activation of scalar alternatives with semantic priming. experiments in linguistic meaning 2. 229–240. 10.3765/elm.2.5371. ronai, eszter & ming xiang. 2024. what could have been said? alternatives and variability in pragmatic inferences. journal of memory and language 136. 104507. rooth, mats. 1992. a theory of focus interpretation. natural language semantics 1(1). 75–116. steffman, jeremy, jennifer cole & stefanie shattuck-hufnagel. 2024. intonational categories and continua in american english rising nuclear tunes. journal of phonetics 104. 101310. van tiel, bob, emiel van miltenburg, natalia zevakhina & bart geurts. 2016. scalar diversity. journal of semantics 33(1). 137–175. 10.1093/jos/ffu017. wagner, michael, elise mcclay & lauren mak. 2013. incomplete answers and the rise-fall-rise contour. in proceedings of the 17th workshop on the semantics and pragmatics of dialogue, 140–149. http://semdial.org/anthology/z13-wagner semdial 0018.pdf. ward, gregory & julia hirschberg. 1985. implicating uncertainty: the pragmatics of fall-rise intonation. language 61. 747–776. 10.2307/414489. watson, duane g, michael k tanenhaus & christine a gunlogson. 2008. interpreting pitch accents in online comprehension: h* vs. l+ h. cognitive science 32(7). 1232–1244. westera, matthijs. 2019. rise-fall-rise as a marker of secondary quds. in secondary content, 376–404. brill. 10.1163/9789004393127 015. yan, mengzhu & sasha calhoun. 2019. priming effects of focus in mandarin chinese. frontiers in psychology 10. 10.3389/fpsyg.2019.01985. 12 proceedings of elm 3: 383-394, 2025 thomas sostarics, eszter ronai, and jennifer cole: relating scalar inference and alternative activation: a view from the rise-fall-rise tune. 394 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms radim lacina & nicole gotzner* abstract. speakers often use scalar words such as warm in a pragmatically strengthened way that results in the conveyed meaning being warm but not hot. in these inferences, known as scalar implicatures, meaning alternatives have been postulated to play a crucial role. upon encountering warm, the informationally stronger alternative hot has been shown to be active in online sentence processing. antonyms (cool) have also been shown to be activated in the same way even though they are standardly assumed to not be involved in scalar implicature derivation. in the current study, we focus on the question of whether both strong scalar alternatives and antonyms are represented in the final mental model of the discourse following scalar implicature derivation. we ran two probe recognition experiments, testing strong scalars and antonyms. we found an interference effect for strong scalars, indicating their representation, but not one for antonyms. thus, we provide evidence that only the strong scalars survive in the eventual representation of the pragmatic meaning of a sentence. keywords. scalar implicatures; alternatives; probe recognition; antonyms; informational strength 1. introduction. in human language, there are many words whose meaning is tied to a particular underlying scale with their contribution being that they describe a certain degree of some property or other (kennedy 2007). often, there are families of related words that differ in the degree of the property which they ascribe. through their semantic relations, they then also give rise to pragmatic inferences communicated by speakers and derived by listeners. take the following example: (1) when i put my foot in, the bath water was warm. when uttering (1), the speaker most likely wishes to convey two meanings at once. firstly, they say that the temperature of the water crossed a certain contextually determined threshold, in other words, that it was at least warm. this would correspond to the literal meaning contributed by the adjective warm (kennedy & mcnally 2005). secondly, there is the added pragmatic inference that while the water was warm, its temperature was not so high as to be considered hot. these inferences have become known as scalar implicatures (horn 1972). we can tell that this other *this research was supported by the german research foundation (dfg) through its emmy-noether programme, a grant awarded to the second author (nr. go 3378/1-1). we would like to thank the audience of elm3 for their valuable comments on this research and stavroula alexandropoulou for collaborating with us on an inter-experimental participant-sharing scheme employed in experiment 1. authors: radim lacina (radim.lacina@uni-osnabrueck.de), institute of cognitive science, osnabrück university; department of czech language, faculty of arts, masaryk university, & nicole gotzner (nicole.gotzner@uni-osnabrueck.de), institute of cognitive science, osnabrück university. author contributions: radim lacina: conceptualisation, methodology, resources, software, data curation, investigation, formal analysis, visualisation, writing – original draft, writing – review and editing; nicole gotzner: conceptualisation, methodology, writing – review and editing, supervision, funding acquisition, project administration. proceedings of elm 3: 201-213, 2025 c©2025 radim lacina and nicole gotzner published by the lsa with permission of the author(s) under a cc by license. 201 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ meaning is in fact pragmatic, because this scalar implicature passes the test of cancellability as seen in the following example (grice 1975): (2) when i put my foot in, the bath water was warm. in fact, it was hot! notice that the conjunction of the two sentences comprising (2) is not a contradiction. what this means is that the negation of the proposition that the water was hot is not entailed by (1). one can clearly see the difference when we substitute warm for its antonym cool. as the reader can ascertain for themself, uttering (3) is at least pragmatically odd: (3) #when i put my foot in, the bath water was cool. in fact, it was hot! in the current study, we are concerned with the eventual representation of these scalar implicatures with special attention paid to the unmentioned stronger scalar alternative hot, which is derived in the case of (1), following an online derivation process. we are also interested in how this contrasts with cases such as (3) where the relationship between two scalar words is that of antonymy. the remainder of this article is structured as follows. we begin with an introduction to the theoretical treatments of scalar implicatures followed by a discussion of the current literature on the online derivation of scalar implicatures and the role of alternative meanings in this process. we then present two probe recognition experiments where we contrast stronger scale-mate alternatives with antonymic ones and show that only the former are retained in the representation of implicated meaning whereas the latter, the antonyms, are not. 1.1. theoretical treatments of scalar implicatures. most accounts of scalar implicatures coming from the fields of formal semantics and pragmatics see alternative meanings as crucial to the derivation of the enriched meaning component (for an overview of various theories, see sauerland 2012). horn (1972) sees the key to implicature derivation in ordered lexical scales. according to this approach, words such as warm and hot or some and all are represented in the lexicon ordered by informational strength. for example, hot is said to be the informationally stronger scale-mate to warm, since they have an asymmetric entailment relationship. when we describe some object as being hot, the truth of that sentence necessitates that the same object is at least warm at the same time. however, when we reverse the expressions, something’s being warm does not mean that it must also be hot. under this view, other scale-mates as well as other related words are irrelevant when it comes to scalar implicature derivation. for example, antonyms such as cool or cold do not play a role in this process and are even not considered to be on the same scale as those adjectives of opposite polarity (see horn 1972; for the split-scale assumption). however, there is some recent evidence pointing in the direction that despite what this assumption might suggest, antonyms could be involved in the derivation of implicatures. peloquin & frank (2016) found that in their computational study of implicatures, including antonyms in their model improved its ability to predict human judgements of pragmatic meaning. skordos & papafragou (2016) found that the presence of the non-entailed alternative none boosted the derivation of some to some but not all implicatures in 5-year-old children. baker et al. (2009) found that when non-entailed alternatives were included in a qud (question-under-discussion) preceding a sentence with a weak implicature-allowing scalar, proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 202 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ this also led to an increase of the acceptance of the implicated meaning. in the case of relative adjectives, of which warm and hot are prime examples (kennedy & mcnally 2005), alexandropoulou & gotzner (2022) found that the pragmatic inferences with these adjectives were supported by the presence of an antonymic contrast in the incremental decision visual world study they ran. absolute adjectives (e.g., breezy) did not show the pattern and people were equally likely to arrive at the implicature meaning with or without this contrast. this suggests that, at least with some types of scales, the cueing the antonymic meaning supports implicature derivation. peloquin & frank (2016) report their computational modelling study of implicature derivation and argue that also including expressions other than just the strong scalar, such as antonyms, improved how well the their model fit the human judgement data. all of this research suggests that alternatives are crucial in the derivation of scalar implicatures, but also that what alternatives exactly are involved and how is still unknown and a matter for current research, albeit with preliminary suggestive evidence that alternatives beyond the informationally stronger ones could also play a role. these questions might in turn be helped by empirical methods that seek to examine what alternatives are present during language processing and when. we review these findings below. 1.2. the activation of scalar alternatives. recently, researchers have started examining the role of alternatives in the online processing of scalar implicatures. similarly to earlier research in the related domain of focus, where it had been shown that alternative meanings are active in the process of comprehension (braun & tagliapietra 2010, husband & ferreira 2016, gotzner et al. 2016), lexical decision experiments were run to test this. de carvalho et al. (2016) ran a masked priming experiment focusing on the lexical association between weak and strong terms such as some and all. what they reported was a pattern of asymmetric priming. weak scalar terms (some) activated their stronger scale mates (all) to a larger degree than the strong ones activated the weak. the researchers interpreted this as evidence for the psychological reality of lexically encoded horn scales (horn 1972). they, however, did not test scalar words in contexts where they could plausibly support implicatures in the minds of comprehenders, since they presented their scalar stimuli as isolated lexical items. ronai & xiang (2023) moved this research forward into the domain of sentence processing and asked whether there is a difference in the activation of the strong term by its weak scale-mate when the latter is embedded in a sentence that could give rise to an implicature as opposed to when presented as an isolated lexical item. they used sentential frames such as the following: (4) zack’s carpet was dirty/patterned. in their sentential experiment, native speakers of american english saw sentences such as (4) either with a related weak scalar prime (dirty) or an unrelated one (patterned). they then completed a lexical decision task on the word filthy, which was the strong scale-mate to dirty. ronai & xiang (2023) found that it was only when the weak scalar primes were embedded within a sentential context that priming occurred. on the other hand, them presenting these words in isolation did not cause participants to react to the strong scalar targets faster. the researchers concluded this to be evidence of their priming effect being indicative of pragmatic-specific activation and argued that strong scalar terms were being involved as meanproceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 203 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ing alternatives in the process of online implicature derivation. a study by lacina et al. (2024) examined two questions regarding the activation of scalar alternatives in processing. they were interested in whether (a) the priming was specific to sentential contexts with a potentially derivable implicature and would disappear when the context would not allow for it and (b) whether alternatives other than just the informationally stronger term were active in the process. firstly, they introduced constituent negation to sentences such as (4) with related primes being not dirty. negation is said to reverse entailment relations, which in the case of (4) results in the presented target word filthy no longer being informationally stronger in this context. consequently, the scalar implicature of dirty, but not filthy is no longer derived. what the researchers found was that in these cases, there was no longer any priming from dirty to filthy. next, they examined whether antonyms were also activated during the process. they replaced the weak scale-mate primes (dirty) with their antonyms (clean) while keeping the same targets (filthy). they found that antonyms activated the targets, similarly to weak scalars. what this research shows is that stronger scalar alternatives are preferentially activated during comprehension in cases where they can support a pragmatic inference and that antonymic meanings are activated in that process at the same time. in the following section, we present an account attempting to capture this pattern. 1.3. the role of alternatives in comprehension. to explain the processing of alternatives in comprehension, the alternative activation account has been posited (husband & ferreira 2016, gotzner 2017). this is view that puts forward a two-step mechanism (see gotzner & lacina (in print) for an application to scalar implicatures). it suggests that lexical alternatives are first activated together with pragmatically irrelevant semantically associated items in early processing by means of domain-general mechanisms of activation spreading. in the following step, the proper alternatives (i.e., the ones relevant to implicature processing) are selected and maintained in further processing. this model is illustrated in a simplified manner in figure 1 below, which takes the example of the scale used in (4) and gives the reader a general idea of what the alternative activation account postulates. the example itself is of the process that begins when a comprehender encounters the weak scalar word dirty within an upward-entailing sentential context, where implicatures are expected to be derived, such as the one in (4). the studies of ronai & xiang (2023) and lacina et al. (2024) have focused on the early stages of processing. they have shown that both strong scalars and antonyms are activated at least during the first stage (step 1 in figure 1). what these studies have not addressed, however, is the state of the comprehenders’ minds once the process of implicature derivation is complete and what is included in the eventual representation of the enriched meaning constructed. in the following sections, we aim to address this gap in knowledge. we first discuss how this question might be approached using the probe recognition task (gernsbacher & jescheniak 1995) and argue for its suitability as a tool to examine representations in the mental models of discourse that comprehenders form. 1.4. the representation of alternatives. above, we reviewed how the lexical decision task, having been first used to study focus alternatives, has then been implemented in the domain of scalar implicatures. in the current section, we introduce another method that was first used to study focus alternatives, but with the goal of studying eventual representations—the probe recognition proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 204 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ initial activation stage selection mechanisms implicature derived antonyms clean stronger scalemates filthy stronger scalemates filthy implicature not filthy step 1 step 2 further processes upon proper alternatives figure 1: diagram illustrating the proposed alternative activation account in the case of encountering the weak scalar word dirty in an implicature allowing context. during step 1, both antonyms and strong scalars are activated. what follows is a selectional phase during which only the stronger scale-mate (filthy) remains activated. next, further processes derived the implicature that is equivalent to the negation of the strong term (not filthy). task (gernsbacher & jescheniak 1995, gotzner et al. 2016). studies on meaning alternatives in the case of focus have found that both contextually mentioned and unmentioned alternatives are represented in the mental model of the discourse (gotzner et al. 2013, 2016, jördens et al. 2020) and that these alternatives are more strongly encoded in memory (fraundorf et al. 2010, norberg & fraundorf 2021, spalek et al. 2014). the method employed in these studies is the probe recognition task, in which participants judge whether a given probe word appeared in a previously presented stimulus. this task has been argued to tap into the eventual strength of representation in the mental model of the discourse (gernsbacher & jescheniak 1995). in this focus work by gotzner et al. (2016), their probe recognition experiments revealed that focus alternatives were reacted to more slowly compared to unrelated words, and that this effect was increased by the presence of focus-sensitive particles such as only or also. their reasoning was that this was due to a competition process between the alternatives and the focused element, making it more difficult to either confirm the presence of mentioned alternatives or reject that of unmentioned ones. such competition could have only occurred in case that these alternatives were actually being represented as being parts of the discourse. mere semantic associates, as opposed to plausible alternatives, did not show an interference effect in the probe recogition task (gotzner proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 205 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ & spalek 2017). thus, the interference effect could be taken as evidence for the representation of a particular alternative. employing this method would, therefore, allow us to answer the question of what scalar alternatives are retained following their initial activation and subsequent pruning during implicature. for those alternatives that we predict to be necessary for this process, namely stronger scale-mates, we should expect an interference effect caused by their competition with the weak scale-mate that is the implicature trigger. other related words, such as antonyms, however, are predicted not to give rise to this interference under both standard theoretical accounts (horn 1972) and the alternative activation account, since they ought to be eliminated following their pruning after an initial activation stage during processing. in other words, they should be absent from the final representation. 2. experiments 1 & 2. in this research, we aim to answer the question of whether scalar alternatives are represented in the mental model of the discourse at a point when implicature derivation is complete. furthermore, we are interested in what kinds of alternatives these are—whether only the informationally stronger alternatives, which are relevant in the derivation, are represented or whether it is also antonyms, which theory considers irrelevant, that are in the final mental model. lacina et al. (2024) found that in earlier processing, antonyms are activated. we therefore ask whether this activation is maintained. we ran two web-based probe recognition experiments, one with weak scale-mate primes and one with antonymic ones in order to answer these questions. our predictions were based on the alternative activation account (gotzner & lacina in print). given that strong scalar terms are said to be relevant for implicature derivation (horn 1972) and that their negation forms a conjunction with the literal meaning of the sentence giving rise to the implicature, we expect the strong term to be included in the eventual representation in the mental model of the discourse. as for antonyms, these do not form any core meaning of the derived implicature under standard theoretical treatments (horn (1972) and following work). while there is evidence that antonyms are active and potentially relevant during the early stages of implicature derivation (see the discussion of alexandropoulou & gotzner 2022, lacina et al. 2024, peloquin & frank 2016; above), they are predicted to not be in the final discourse representation. as argued for above, we operationalise the strengthened representation of a particular element in the mental model of the discourse by its interference effect in the probe recognition task. we expect the strong scalar targets to be rejected slower when preceded by their weaker scale-mate as opposed to by an unrelated word. the same effect should not occur with antonymic primes. 3. methods. 3.1. data availability. our experimental stimuli as well as the statistical analysis script are freely available for downloading on the open science framework platform. they may be reached using the following hyperlink: https://osf.io/khfgb/. 3.2. participants. we recruited 79 participants for experiment 1 and 80 for experiment 2 (159 in total) on the prolific platform. they were monolingual native speakers of american english, born in the us and american nationals, between the ages of 18 and 35, and participants with a 100 approval rate on prolific. experiment 1 was run together with another web-based eye-tracking proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 206 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ experiment not reported here. participants were invited to take part in the eye-tracking experiment first. in case they failed certain web-camera calibration checks, they were redirected to experiment 1, here discussed. participants for experiment 2 were recruited directly. the demographics were the following. in experiment 1, 45 women, 32 men and one person of diverse gender took part. one participant chose the option of preferring to not give their gender information. in experiment 2, these numbers were 35 women, 42 men, two people of diverse gender, and one person did not disclose their gender. the mean age of participants in experiment 1 was 27.9 (sd = 4.8) and in experiment 2, it was 28.1 (sd = 4.3). 3.3. materials. we used the 60 items implemented in experiment 3 of ronai & xiang (2023) and experiment 2 of lacina et al. (2024). take the following example item: (5) zack’s carpet was dirty. [related weak scale-mate, experiment 1] (6) zack’s carpet was clean. [related antonym, experiment 2] (7) zack’s carpet was patterned. [unrelated, experiments 1 and 2] the target word that participants reacted to was the same across experiments and conditions and was always the strong term relative to the weak scale-mate prime, in this case filthy. each experiment contained one of the two related items, weak scale-mates in experiment 1 and antonyms in experiment 2. both experiments shared the same unrelated items. this means that the correct response in the probe recognition task was always no in the experimental items, since the target word did not appear in the stimulus. additionally, we created 60 fillers, where the correct answer was yes. these fillers were again taken from the study of ronai & xiang (2023). our modification consisted in changing the probes from non-words to words appearing in the sentence part of the filler in order to fit the different requirement of the probe recognition task. 3.4. procedure. when participants arrived at the website with the experiments, either directly from prolific in the case of experiment 2 or via redirection from another experiment in the case of experiment 1, they first read a consent form. after this, they received the instructions for the experiment. their task was to read sentences in the rapid serial visual presentation mode (rsvp) followed by probe words. they were asked to indicate whether a given probe word appeared anywhere in the preceding sentence. they were instructed to use j for yes and f for no. following practice items, the experiment itself commenced. sentences presented in the rsvp mode (potter 2018) were displayed at the rate of 350ms per word. after the last word of each sentence, 2000ms of a blank screen elapsed. after this, the probe word associated with the stimulus sentence appeared. overall, the participants rated 60 experimental items and 60 filler items. the ratio of the correct yes and no responses was 1:1. items were distributed based on the latin square design. 3.5. analytical steps. we first excluded all trials of those participants whose accuracy on the combined set of experimental and filler items was below 90%. in experiment 1, this amounted to four participants. in experiment 2, we excluded two according to the same criterion. we then took all the trials of the remaining participants and filtered out all the trials with incorrect responses. the response times of these trials were then log-transformed for the analysis. we ran three linear mixed effects models. in all models, we included the maximal random effects structure allowed appropriate for the data that converged (barr et al. 2013). there was one model proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 207 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ per experiment, where we included the single fixed effect of relatedness (related or unrelated) that was treatment coded with the related condition being 1 and the unrelated 0. the final model was the joint analysis one, where we first created a combined dataset and added the factor of experiment with sum coded conditions of weak scalar (as 1, experiment 1) and antonym (as -1, experiment 2). in this analysis, the condition of relatedness was sum coded too with related being 1 and unrelated -1. we included the fixed effects of the two factors and their interaction. here, in the random effects structure, we did not attempt a model with random slopes for the effect of experiment or its interaction with relatedness for participants as this was a betweensubject factor. 4. results. the reader may consult the combined graphical representation of the response time results of experiments 1 and 2 in figure 2. below, we first give the results of the individual models for each experiment followed by the joint analysis. 738.2 728.1 765.9 741.2 730 740 750 760 770 related unrelated condition r t ( m s) exp 2 (antonyms) exp 1 (weak scalars) figure 2: mean reaction times in ms by condition with associated standard errors in experiments 1 and 2 4.1. experiment 1: weak scalars. the final model that converged had a random structure containing random intercepts for participants and items as well as random slopes for the effect proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 208 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ of relatedness for both. the model revealed that a significant effect of relatedness was present (β = 0.0337, se = 0.0099, df = 52.11, t = 3.416, p = 0.00124**). this effect was in the positive direction, meaning that related weak scalar primes were associated with longer response times to the targets compared to when these targets were preceded by sentences with unrelated critical words. 4.2. experiment 2: antonyms. the model run on the data from experiment 2 included random intercepts for participants and items as well as random slopes for the effect of relatedness for participants. as for the single effect of relatedness, this proved insignificant (β = 0.0087, se = 0.0076, df = 75.78, t = 1.144, p = 0.256). related antonymic primes, therefore, were not found to cause any differences in the rejection of the targets compared to unrelated ones. 4.3. joint analysis. the random effects structure of the joint analysis model that ended up converging without errors included random intercepts for participants and items as well as random slopes for the main effect of relatedness for both. the main effect of experiment was insignificant (β = 0.0126, se = 0.0170, df = 151, t = 0.738, p = 0.46155). there was a significant main effect of relatedness: β = 0.0106, se = 0.0031, df = 58.7, t = 3.418, p = 0.00115**. crucially, the interaction between relatedness and experiment was significant: β = 0.0061, se = 0.0027, df = 146.3, t = 2.260, p = 0.02528*. 5. general discussion. we investigated the representation of alternatives in the mental model of the discourse following the comprehension of scalar implicature-allowing sentences. we asked whether both strong scalars and antonyms were represented therein and ran two probe recognition experiments in this pursuit. based on the alternative activation account (gotzner 2017, gotzner & lacina in print), we predicted that strong scalars would be present and antonyms would be absent. previous studies found that stronger scale-mates, which have been proposed to be necessary for implicature computation in the theoretical literature (horn 1972, sauerland 2012), are in fact activated in early stages of processing following exposure to a scalar word within an implicatureallowing sentence (ronai & xiang 2023). further, it was found that even antonyms, which are standardly assumed not to be involved in the implicature derivation process, are activated at the same time (lacina et al. 2024). the results of our current probe recognition study showed that strong scalar items (filthy) were more strongly represented in the model when their weaker scale-mates were present in the preceding stimulus as opposed to when the prime was an unrelated word (experiment 1). this was evidenced by an interference effect reflected in increased response times. antonymic primes (clean), on the other hand, did not lead to an increased strength of representation for the same targets (experiment 2). there, we saw no difference with the semantically unrelated condition. our joint analysis then showed that this difference was present when the two datasets were directly compared—the type of related prime, i.e., weak scalar or antonym, influenced the size of the interference effect with weak-scalar primes associated with higher response times. what this pattern shows is, we argue, that the informational strength relations between what is encountered in the course of comprehension and potential alternatives has an impact on what ends up being present in the representation of the combined literal and pragmatic meaning in real-time proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 209 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ comprehension. strong scalar terms such as hot or filthy are relevant for the representation of the enriched meaning gained by implicature derivation, whereas antonyms (cool, clean) are not. this interpretation is based on the ideas presented in gotzner et al. (2016), which examined the influence of the particle only on focus alternatives and the lack of this effect with mere associates reported in gotzner & spalek (2017). the suggestion is that those elements that are alternatives show an interference effect, since they are in competition with the element present in the sentence. these are then harder to distinguish from each other. the reasoning is the same in the scalar implicature case. given that we assume that the final representation of sentences with scalar words such as my soup was warm includes the negated strong term hot, which could be said to be roughly equivalent to the representation of my soup was warm but not hot, we expect comprehenders to experience difficulty in indicating that hot was not in fact present in the sentence they had just read. the present results relate to those reported in the study of lacina et al. (2024), who found that a the point of 650ms after encountering a scalar words, antonyms are still activated. their data could not disentangle two competing options—antonyms being selected for and used in implicature derivation or only activated due to the first step of domain-general activation spreading in the lexical-semantic network. what the current data add here is evidence in favour of the latter, namely that antonyms are not directly involved in the process of implicature derivation and are only the remnants of the first activation step. thus, our data are in line with the predictions of the alternative activation account (see figure 1 for a schematic illustration). what our study targeted was the stage at which the implicature derivation processed is finished and antonyms have already been eliminated during stage 2. one potential criticism of our conclusions might be that what our results reflect is simply the difference in similarity between the weak scalars and the targets and the antonyms and the same targets. it could be that antonyms are, on average, less related to the targets compared to the weak scalar primes, causing a difference to appear that is not related to any pragmatic processes, but a much lower-level contrast between the primes. in order to deal with this issue, we conducted an additional analysis where we included similarity scores to test whether our results could be put down to this issue. firstly, we took the similarity scores between the related (weak scalar or antonym) primes and the targets as well as the scores between the same targets and the unrelated primes as reported in lacina et al. (2024). we ran a post-hoc analysis where we included the similarity scores between the prime and target as a fixed effect reported below. in order to explore the influence of semantic similarity on our results and whether there was any difference in its effect on the priming caused by antonymic and weak scalar primes, we ran a nested effects model on the related conditions subset of the data. we nested the effect of similarity within the experiment factor (i.e., within weak scalars or antonyms). the model included random intercepts for participants and items and the random slope for similarity in the case of participants. there was no main effect of experiment (i.e., type of related prime): β = −0.0216, se = 0.0280, df = 141.9311, t = −0.771, p = 0.4422. as for the effect of similarity within weak scalar and antonym primes, the model revealed that it only had an effect within experiment 1 where weak scalars were the primes (β = 0.1380, se = 0.0627, df = 66.3464, t = 2.203, p = 0.0311*). this effect was insignificant within experiment 2 (β = 0.0042, se = 0.0536, df = 65.9823, t = 0.078, p = 0.9382). proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 210 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ what this suggests is that the gradient effect of semantic similarity was only in place when the relationship between the prime and target was that of the weaker and stronger scale-mate. this did not occur when the relationship in question was that of antonymy. we might interpret this result as suggesting that in the case of weak-scalar primes, the strong scalar targets are in fact being represented in the mental model of the discourse and therefore, comprehenders experience difficulty rejecting them in the probe recognition task. the effect of similarity can then be understood conditionally: when the element is represented, the degree of difficulty is then modulated by semantic closeness with those more associated being harder to reject. in the case of antonymic primes, however, the targets are presumably not represented in the model, since they are not required for any pragmatic derivations. here, similarity has no effect, since there is no competition due to representation that could give rise to a slow-down effect. this is also in line with the main result of experiment 2, which found that antonymic primes did not differ from unrelated ones in their influence on the rejection of target words. 6. conclusion. this study examined the strength of representation of scalar terms in the mental model of the discourse when these were preceded either by their weak scale-mates or by their antonyms. we found an interference effect, indicative of increased representation strength, only for the weak scalar primes (experiment 1) and not antonymic ones (experiment 2). the interpretation here pursued is that this is due to the fact that only strong scalar terms maintained in the final product of implicature derivation, as per most standard theories (sauerland 2012), yet antonyms are irrelevant and are therefore not preferentially represented. this is in line with the alternative activation account (gotzner 2017, gotzner & lacina in print) which proposes an initial activation phase in which both strong scalars and antonyms are present with a subsequent narrowing to only the relevant alternatives. references alexandropoulou, stavroula & nicole gotzner. 2022. negation, polarity, scale structure: different inferences of absolute adjectives. in proceedings of sinn und bedeutung, vol. 26, 35–54. baker, rachel, ryan doran, yaron mcnabb, meredith larson & gregory ward. 2009. on the non-unified nature of scalar implicature: an empirical investigation. international review of pragmatics 1(2). 211–248. barr, dale j, roger levy, christoph scheepers & harry j tily. 2013. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language 68(3). 255–278. braun, bettina & lara tagliapietra. 2010. the role of contrastive intonation contours in the retrieval of contextual alternatives. language and cognitive processes 25(7-9). 1024–1043. de carvalho, alex, anne c reboul, jean-baptiste van der henst, anne cheylus & tatjana nazir. 2016. scalar implicatures: the psychological reality of scales. frontiers in psychology 7. 1500. fraundorf, scott h, duane g watson & aaron s benjamin. 2010. recognition memory reveals just how contrastive contrastive accenting really is. journal of memory and language 63(3). 367–386. proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 211 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ gernsbacher, morton ann & jörg d jescheniak. 1995. cataphoric devices in spoken discourse. cognitive psychology 29(1). 24–58. gotzner, nicole. 2017. alternative sets in language processing: how focus alternatives are represented in the mind. springer. gotzner, nicole & radim lacina. in print. generating and selecting alternatives for scalar implicature computation: the alternative activation account and other theories. in alternatives in grammar and cognition, palgrave macmillan. gotzner, nicole & katharina spalek. 2017. role of contrastive and noncontrastive associates in the interpretation of focus particles. discourse processes 54(8). 638–654. gotzner, nicole, katharina spalek & isabell wartenburger. 2013. how pitch accents and focus particles affect the recognition of contextual alternatives. in proceedings of the annual meeting of the cognitive science society, vol. 35, . gotzner, nicole, isabell wartenburger & katharina spalek. 2016. the impact of focus particles on the recognition and rejection of contrastive alternatives. language and cognition 8(1). 59–95. grice, herbert p. 1975. logic and conversation. in speech acts, 41–58. brill. horn, laurence robert. 1972. on the semantic properties of logical operators in english. university of california, los angeles. husband, e matthew & fernanda ferreira. 2016. the role of selection in the comprehension of focus alternatives. language, cognition and neuroscience 31(2). 217–235. jördens, kim, nicole gotzner & katharina spalek. 2020. the role of non-categorical relations in establishing focus alternative sets. language and cognition 12(4). 729–754. kennedy, christopher. 2007. vagueness and grammar: the semantics of relative and absolute gradable adjectives. linguistics and philosophy 30(1). 1–45. kennedy, christopher & louise mcnally. 2005. scale structure, degree modification, and the semantics of gradable predicates. language 345–381. lacina, radim, stavroula alexandropoulou, eszter ronai & nicole gotzner. 2024. scalar alternative activation in implicature processing: a lexical decision study with antonyms and negation. psyarxiv 10.31234/osf.io/r3q79. norberg, kole a & scott h fraundorf. 2021. memory benefits from contrastive focus truly require focus: evidence from clefts and connectives. language, cognition and neuroscience 36(8). 1010–1037. peloquin, benjamin n & mike frank. 2016. determining the alternatives for scalar implicature. in proceedings of the annual meeting of the cognitive science society, vol. 38, . potter, mary c. 2018. rapid serial visual presentation (rsvp): a method for studying language processing. in new methods in reading comprehension research, 91–118. routledge. ronai, eszter & ming xiang. 2023. tracking the activation of scalar alternatives with semantic priming. in experiments in linguistic meaning, vol. 2, 229–240. https://doi.org/10.3765/elm.2.5371. sauerland, uli. 2012. the computation of scalar implicatures: pragmatic, lexical or grammatical? language and linguistics compass 6(1). 36–49. skordos, dimitrios & anna papafragou. 2016. children’s derivation of scalar implicatures: alternatives and relevance. cognition 153. 6–18. proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 212 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ spalek, katharina, nicole gotzner & isabell wartenburger. 2014. not only the apples: focus sensitive particles improve memory for information-structural alternatives. journal of memory and language 70. 68–84. proceedings of elm 3: 201-213, 2025 radim lacina and nicole gotzner: only the (informationally) stronger survive: a probe recognition study with scale-mates and antonyms. 213 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ‘negation-blind’ n400 effect disappears when lexical priming is controlled daiki asami, chao han, jacob burger, deanna dunlop, yue lu, effah yahya m morad, chenyue zhao, & arild hestvik∗ abstract. previous erp studies showed that false affirmative sentences elicited a larger n400 than their true versions, but they found the reverse pattern when the sentences were of negative form as if n400 was blind to negation. this negationblind n400 pattern arguably constituted evidence for two-step accounts of negation processing: when processing negative sentences, a comprehender first computes an internal proposition and then considers the negation. however, the prior studies were confounded by a lexical priming relation between subject and object. therefore, it was an open question whether or not the observed erp pattern really reflected the twostep process. to tackle this question, we conducted an erp experiment, using sizecomparison statements where subjects and objects are semantically unrelated. this design allowed us to remove the priming confound. we predicted that if the previous negation-blind n400 pattern is unrelated to lexical priming, it would be replicated; if not, it would disappear. the result was consistent with the second prediction. this suggests that the previously observed negation-blind n400 pattern does not necessarily constitute evidence for two-step accounts of negation processing. keywords. sentence processing; negation; n400; two-step accounts; lexical priming 1. introduction. processing of negative sentences (e.g., a whale is not a fish.) has attracted much attention in psycholinguistics, as evidenced by many reviews (kaup et al. 2007, tian & breheny 2019, christensen 2020, kaup & dudschig 2020, papeo & de vega 2020, dudschig et al. 2021). these reviews showcase various theories of negation processing. among them, two-step accounts (clark & chase 1972, carpenter & just 1975, kaup et al. 2006, 2007, kaup & dudschig 2007, palaz et al. 2020) postulate that negation processing involves two steps, as schematized in (1): a comprehender first computes a to-be-negated internal proposition and then considers negation.1 (1) step 1: ⟦a whale is a fish⟧ step 2: ⟦not⟧ (⟦a whale is a fish⟧) two-step accounts received neurological support from event-related potential (erp) research, that is, a ‘negation-blind’ pattern of n400 results (fischler et al. 1983, kounios & holcomb 1992, ∗we thank xu qing for the conception of the experiment and neemias silva de souza filho for helpful feedback. we are also grateful to the audience at elm3 for their comments. authors: daiki asami, university of delaware (daiasami@udel.edu), chao han, university of toronto scarborough (chao.han@utoronto.ca), jacob burger, university of delaware (jburger@udel.edu), deanna dunlop, gallaudet university (deanna.dunlop@gallaudet.edu), yue lu, university of delaware (sunnylu@udel.edu), effah yahya m morad, university of delaware (emorad@udel.edu), chenyue zhao, temple university (chenyue.zhao@temple.edu), arild hestvik, university of delaware (hestvik@udel.edu). this study was conducted while all the authors were affiliated with the university of delaware. 1we are agnostic about exact representations of the proposition to which not is applied. for this reason, we use double square brackets, following the formal semantic tradition. this does not undermine the main claims made by the previous studies. proceedings of elm 3: 19-31, 2025 c©2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al. published by the lsa with permission of the author(s) under a cc by license. 19 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ dudschig et al. 2019, haase et al. 2019, palaz et al. 2020).2 an n400 is a negative erp component with a peak at approximately 400 ms after the onset of every content word; its amplitude is sensitive to a wide range of semantic factors (see kutas & federmeier (2011) for a review). among them, content words that cause world knowledge violations increase the n400 amplitude (hagoort et al. 2004, hald et al. 2007, metzner et al. 2015, nieuwland et al. 2019). consistent with this, prior studies found that false affirmatives in (2b) elicited a larger n400 relative to their true counterparts in (2a) (fischler et al. 1983, kounios & holcomb 1992, dudschig et al. 2019, palaz et al. 2020).3 (2) a. a robin is a bird. (true affirmative) b. a robin is a tree. (false affirmative) (fischler et al. 1983: 402) crucially, the reverse pattern was observed in the negative sentences: the true sentences in (3a) elicited a larger n400 than their false versions in (3b), as if the n400 was only sensitive to the to-be-negated proposition and ‘blind’ to negation. (3) a. a robin is not a tree. (true negative) b. a robin is not a bird. (false negative) (fischler et al. 1983: 402) this n400 pattern falls into place under two-step accounts: when comprehending the true negative sentence (3a), only the false internal proposition (i.e., a robin is a tree) is computed in the n400 time window; consequently, it elicits a larger n400 than the false negative sentence (3b), in which the core proposition (i.e., a robin is a bird) is true. however, as acknowledged by the original negation-blind n400 study (fischler et al. 1983: 407–408) and later brought up by others (nieuwland & kuperberg 2008, wiswede et al. 2013, haase et al. 2019, he et al. 2022), the previous erp studies were confounded by a lexical priming relation between subject and object, which is known to be inversely correlated with n400 amplitude (kutas & hillyard 1984, bentin et al. 1985, rugg 1985). for instance, the target word nurse leads to a reduced n400 when it follows the semantically related word doctor compared to when it follows a semantically unrelated word cat. this is because the prime word doctor not only activates the representation of itself but also pre-activates related words or concepts within the lexicon; consequently, pre-activated words require less cognitive costs than non-preactivated words, resulting in a smaller n400 amplitude. returning to our previous discussion on erp studies, the true affirmative sentences exhibited the semantic relatedness between subject and object (e.g., robin/bird in (2a)) while their false counterparts did not (e.g., robin/tree in (2b)). the same contrast was also true of the comparison of the false and true negatives (e.g., robin/bird and robin/tree in (3b) and (3a), respectively). as a result, the sentences with the related pair of words (i.e., (2a) and (3b)) may attenuate the n400, relative to those with the unrelated pair of words (i.e., (2b) and (3a)). if this is correct, we may interpret the previously observed negation-blind n400 effect as a reflection of the lexical priming. to directly test this alternative interpretation, haase et al. (2019) conducted an erp experiment using sentences such as (4).4 they intended to control the semantic relatedness by increasing the 2contrary to these studies, nieuwland & kuperberg (2008) found a ‘negation-sensitive’ n400 pattern such that false sentences elicited a larger n400 in both affirmative and negative sentences. see section 4 for the relevant discussion. 3throughout this paper, we underline the critical chunk after which the erps are averaged. 4the original stimuli were in german but we present only the english translations. proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 20 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ sentence form truth affirmative negative true (a) a tiger is bigger than a guitar. (b) a mouse is not bigger than a guitar. false (c) a tiger is smaller than a guitar. (d) a mouse is not smaller than a guitar. table 1: example stimuli in the four conditions, crossing the two truth values (true vs. false) and the two sentence forms (affirmative vs. negative). semantic overlap between subject and object. their experimental sentences used celebrity names (e.g., george clooney) and profession nouns (e.g., actor) for subjects and objects, respectively. (4) a. george clooney currently is an actor in the usa. (true affirmative) b. george clooney currently is a singer in the usa. (false affirmative) c. george clooney is not a singer in the usa. (true negative) d. george clooney is not an actor in the usa. (false negative) (haase et al. 2019: 1) they reasoned that the use of two profession nouns from the related field (e.g., singer and actor) increased the semantic relatedness between subject and object, thereby removing the priming confound. their results revealed an n400 effect in the comparisons of false and true affirmatives ((4b) vs. (4a)) as well as true and false negatives ((4c) vs. (4d)), although the latter comparison was not statistically significant. haase et al. (2019) took the non-significant negation-blind n400 effect as partial, if not strong, support for two-step accounts of negation processing. however, their design had an issue because a celebrity name is more related to one occupation than the other. for instance, george clooney is more semantically associated with actor than singer. thus, their experiment was still confounded by the lexical priming. therefore, it was still open whether the previously observed negation-blind n400 results stemmed from a process assumed by two-step accounts or lexical priming. to tackle this issue, we conducted an erp experiment using size comparison statements (table 1) where the subject and object differed from each other in terms of animacy and semantic category, regardless of truth values and sentence forms. this design allowed us to remove the priming confound by eliminating the semantic relatedness, rather than increasing it, unlike haase et al. (2019). we made two predictions, as in (5) and (6). the first prediction was that if the lexical priming was weakly associated with the previously observed erp pattern, we would observe a larger truthsensitive n400 in a word that makes an affirmative sentence false, compared to the same word that makes it true. additionally, if the two-step negation processing is psychologically real, we would see a negation-blind n400 effect in the opposite subtraction (5b). the second prediction was that if the lexical priming was the source of the previous negation-blind n400 pattern, the critical words would elicit roughly equal n400 amplitude regardless of truth values and sentence forms; as a result, we would observe no n400 effect in the comparisons of interest (i.e., (6a) and (6b)). (5) a. false affirmative − true affirmative = n400 effect proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 21 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ b. true negative − false negative = n400 effect (6) a. false affirmative − true affirmative = no n400 effect b. true negative − false negative = no n400 effect 2. methods. 2.1. participants. 30 undergraduate students (27 females; 3 males) participated and were compensated with extra credit. all were native speakers of english, had normal or corrected to normal vision, and reported no history of neurological diseases or medication. they gave written informed consent before data collection. the study was approved by the irb of the university of delaware. 2.2. design and stimuli. the current experiment had a 2 × 2 within-subjects design. the first factor—truth value—was whether a sentence was true or false (true vs. false). the second factor—sentence form—was whether the sentence included not or not (affirmative vs. negative). crossing the two factors yielded four conditions (table 1). experimental stimuli were comparative constructions describing a size comparison with is (not) {bigger/smaller} than as a predicate. we created 40 sentences for each condition (160 sentences in total). we also created 10 sets of fillers for each condition (40 sentences in total). half of them used five adjectives whose first letter was b (brighter, busier, braver, broader, and bumpier); the other half used those beginning with s (saltier, sweeter, scarier, sicker, and softer). these adjectives exhibit partial similarity to bigger and smaller in the experimental items in terms of orthography and phonology. we utilized this similarity to prevent participants from developing semantic satiation (e.g., (not) bigger and (not) smaller become less meaningful as a function of repetition) (jakobovits & lambert 1962, smith & klein 1990) or adopting any heuristic strategy (e.g., mapping not bigger into smaller without computing the compositional meaning). each participant saw all 200 sentences in a single experiment. however, we expected that results would be confounded by the facilitated processing due to the repetition priming if sentences containing the same words (e.g., tiger) appeared multiple times within a short interval. to rule out this confound, we distributed 200 sentences (160 experimental sentences and 40 filler sentences) into four lists using a latin square design. we then randomly assigned these lists to four individual blocks in each experiment so that the participants never saw more than one item from the same set within each block. 2.3. procedure. our experiment employed a speeded truth-value judgment task (figure 1). to begin each trial, the participant was prompted to press 1 or 5 on the response box. upon the button press, a fixation element appeared for 1,000–1,400 ms, and then a sentence was presented with a rapid serial visual presentation paradigm. the sentence was divided into four chunks (e.g., a tiger / is (not) / bigger than / a guitar.). each chunk was presented for 175 ms with an 800 ms interstimulus interval. upon the presentation of the final chunk, the participant judged whether the sentence was true or false by pressing 1 (= true) or 5 (= false) as quickly and accurately as possible. the allotted response time was 4,000 ms. reaction time and accuracy of the response were recorded. upon the truth-value judgment, the feedback element appeared for 1,000 ms. if the participant did not respond in 4,000 ms, the feedback no response was displayed and the response was recorded as a missing value. the procedure was repeated 50 times in four blocks (200 trials in total). the proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 22 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ press 1 to begin the sentence. 1 5 175 ms + 800 ms isia tiger is not bigger than a guitar. 1 5 1 = true; 5 = false correct! 95.00 average percent correct ++++ 1,000 – 1,400 ms 1,000 ms figure 1: trial structure. orders of the blocks and the items were randomized for each participant. the participants had eight practice trials before the actual experiment. 2.4. eeg recording and data processing. we recorded eeg using 64-channel hydrocel geodesic sensor nets with a sampling rate of 250 hz with cz as the reference electrode. the continuous eeg data were run through a 0.1 hz two-pass fir high-pass filter, followed by a 40 hz two-pass butterworth low-pass filter. we segmented the data into four conditions (true affirmative, false affirmative, true negative, and false negative) excluding those with inaccurate responses. each segment was time-locked to the onset of the object chunk and contained a 200 ms pre-stimulus period for baseline correction followed by a 1,000 ms post-stimulus period. to remove common artifacts (eye blinks, muscle movements, and saccade movements), the eeg data underwent the automated artifact correction with the multi-algorithm artifact correction (maac) procedure (dien 2024). for eye blinks and saccades, each participant’s data underwent independent component analysis (ica) decomposition, and components that correlated at r = 0.95 with predefined blink and saccade templates were subtracted from the data, before the data were reconstituted by remixing the components. bad channels were replaced with spline interpolations of the surrounding channels. after this step, the data were baseline corrected again and re-referenced to the average of all channels. each participant’s single trial data was then averaged into the four conditions. 2.5. erp analysis plan. the aim of the experiment was to identify the brain response to true vs. false sentences and whether this response materialized as an n400 effect. the n400 is typically described as a centro-parietal negativity in the 300–500 ms time window. one approach to analysis would then be to simply average cz, pz, or some collection of centro-parietal midline electrodes and the 300–500 ms time window for each condition and participant and submit the resulting voltage values to an anova. however, each experiment is slightly different, and the erp may jitter temporally and/or spatially depending on the particular task demands, stimulus presentation modes, and participant populations. a priori defined time windows and electrode clusters may therefore not fit the observed data precisely. to improve the precision of time windows and electrode region selection relative to the data set at hand, we therefore utilized sequential temporospatial principle component analysis (pca; dien 2010, 2012, dien & frishkoff 2005), which decomposes the observed surface mixture data into underlying temporal and spatial components of the brain response to the experimental manipulations. specifically, we used the two false minus proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 23 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 2: mean accuracy (right) and rt (left) by condition. true difference waves for affirmatives and negatives as input to the temporal pc, because the experiment sought to identify whether the n400 differed for the two truth value conditions. the pca decomposes the 0–1,000 ms segment into latent time events reflecting the differential effect of truth value differences. we then spatially decompose each temporal factor with ica to determine the main electrode source of those temporal events. to assess the statistical significance of condition differences in the observed pca components, we planned to use the resulting factor scores in each latent time/space component and each cell and participant as dependent measures, and analyze whether the affirmative and negative difference scores differed significantly from zero (a main effect of truth value), and from each other (an interaction with sentence form). 3. results. 3.1. behavioral results. we conducted a 2 x 2 repeated measures anova with sentence form and truth value as independent variables. the right panel in figure 2 shows the mean accuracy for each condition. for accuracy, we found a significant main effect of sentence form such that affirmatives were judged more accurately than negatives (91 vs. 79%; f(1,29) = 93.5, p < 0.001). we also found the significant main effect of truth value: false sentences were judged more accurately than true sentences (83 vs. 88%; f(1,29) = 11.1, p < 0.01). finally, there was a sentence form by truth value interaction such that the difference between true and false affirmatives (91 vs 92%) was smaller than the difference between true and false negatives (76 vs. 83%; f(1,29) = 8.2, p < 0.01). the left panel in figure 2 shows the mean rt by condition. for rts, we used data that evoked correct responses. we found the significant main effects of sentence form and truth value: affirmatives were faster to judge than negatives (1170 vs. 1516 ms; f(1,29) = 131.5, p < 0.001) and true sentences were faster to judge than false ones (1317 vs. 1368 ms; f(1,29) = 4.59, p < 0.05). 3.2. erp results. proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 24 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 3: waveforms (left) of microvolt-scaled factor loadings from the temporo-spatial factor tf2sf1 and the topographic map (right) of the difference between the two false-true differences: the false-true difference in the negative form minus the false-true difference in the affirmative form. 3.2.1. temporo-spatial pca and ica. to analyze our erp data without experimenter bias, we conducted a temporo-spatial pca/ica of all data (using all time samples and all electrodes) with the erp pca tool matlab toolbox (dien 2010). this exploratory step would determine the latent effects, if any, of truth-value computation on the eeg activity. we followed the recommended analysis pipeline in the erp pca toolkit tutorial (dien 2010), using temporal pca based on the covariance of time samples. this was followed by a spatial ica for each temporal factor based on the covariance of electrodes, to further separate out distinct spatial sources of variance within each temporal factor. to determine if the data contained a response to the difference between true and false sentences, we created a main effect of truth difference wave (i.e., (false affirmative + false negative) − (true affirmative + true negative) for each participant). this single dependent measure was then used as input for the pca. to reduce dimensionaltiy in the time domain, we retained 13 factors from the temporal pca solution, which determines 93% of the total variance. following the recommendations in dien (2010), we limit analysis to only factors that account for a sizable amount of variance, arbitrarily set to more than 6% of the total variance. this was true for the first four temporal factors (tfs): tf1 (peaked at 880 ms and accounted for 34% of the variance), tf2 (576 ms, 20%), tf3 (384 ms, 10%), and tf4 (728 ms, 7%). we next applied a spatial ica decomposition to each of the temporal factors, to further narrow down the different spatial regions in each temporal factor according to how much variance is accounted for. four spatial factors were retained for each temporal factor. temporal factor 2, spatial factor 1 (f2sf1) exhibited a left anterior negativity, and was thus taken as a latent factor reflecting an lan effect in the grand average voltage data. figure 3 shows that the effect was carried by the negated sentence. in contrast to a clear lan effect, our visual inspection of the temporal and temporo-spatial factors revealed no n400 effect. for an inferential test, we reconstructed tf2sf1 in voltage space by multiplying the factor loadings with the factor scores (dien & frishkoff 2005), resulting in two reconstructed difference waves: affirmative (fa − ta) and negative (fn − tn). we then obtained the dependent measure by averaging the amplitudes of each difference wave over the pca-delimited time window (506–668 ms) and 11 frontal electrodes (e5, e6, e8, e9, e10, e11, e12, e17, e55, proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 25 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ e58, e63). we tested for the presence of a lan effect in the affirmative and negative sentence forms separately by conducting a one-sample t-test for each sentence form to evaluate whether the difference in amplitudes between the false and true conditions was significantly different from 0. additionally, we performed a paired-sample t-test to compare the two false-true difference amplitudes between the two sentence forms, assessing whether the lan effect differed across the sentence forms. the results showed that the effect was not significant for the affirmative condition (t(29) = -0.17, p = .885) but was significant for the negative condition (t(29) = -2.33, p = .027). in addition, the amplitude difference between the two conditions was significant (t(29) = 2.53, p = .017). the inferential results confirmed a lan effect driven by the false-true difference in the negative (fn − tn) condition (figure 4). figure 4: boxplot of reconstructed tf2sf1 voltage data for the affirmative and negative conditions. each dot represents an individual data point. figure 5 shows the waveform of the ica regionalized mean channel for the main effect of truth value in the affirmative (left) and negative (right) conditions. figure 5: waveform plots with 84% confidence intervals. 4. discussion. our behavioral results replicated prior findings (clark & chase 1972, carpenter & just 1975, fischler et al. 1983, palaz et al. 2020), indicating that our experiment tapped into a similar cognitive process as the previous studies. however, the erp results revealed no n400 proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 26 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ effect in the relevant comparisons (i.e., false affirmative − true affirmative or true negative − false negative). this erp result was consistent with our second prediction: if the previously observed erp pattern came from the lexical priming, the n400 effect would not be observed. instead of the n400 effect, we found the main effects of truth value and sentence form, and their interaction in a lan effect during the 504–673 ms time window. although we must leave a precise explanation to future research, we suspect that this lan effect has to do with the truth verification process. no n400 effect suggests that the final word in the experimental sentences elicited a roughly equal n400 amplitude across conditions. we attribute this result to the lack of semantic relation between subject and object in our stimuli. if this is correct, the previously observed negation-blind n400 effect can be said to result from the lexical priming: the critical word attenuated the n400 when preceded by a semantically related word, compared to being preceded by an unrelated word. this lexical priming account can explain findings by some erp studies on negation processing (fischler et al. 1983, kounios & holcomb 1992, dudschig et al. 2019, haase et al. 2019) but it is silent about the pragmatics-related negation processing reported in nieuwland & kuperberg (2008) and palaz et al. (2020). hence, let us discuss these two studies in detail below. contrary to other erp studies, nieuwland & kuperberg (2008) found a ‘negation-sensitive’ n400 pattern when negative sentences appeared under pragmatically appropriate contexts. specifically, they found that false negative sentences elicited a larger n400 than their true versions, when preceded by contextual phrases (e.g., (7b) vs. (7a), where with proper equipment serves as a contextual phrase). (7) a. with proper equipment, scuba-diving isn’t very dangerous ... (true negative) b. with proper equipment, scuba-diving isn’t very safe ... (false negative) (nieuwland & kuperberg 2008: 1214) our lexical priming account is silent about this negation-sensitive n400 pattern because it is purely concerned with a low-level lexical process. however, this does not mean that it is flawed. previous erp studies showed that a high-level pragmatic process can override the low-level lexical process (nieuwland & van berkum 2006, nieuwland et al. 2010, filik & leuthold 2008, hunt iii et al. 2013).5 specifically, words that caused semantic anomaly did not elicit a large n400 when embedded under pragmatically appropriate contexts. for instance, nieuwland & van berkum (2006) found that salted caused a larger n400 than in love in the sentence the peanut was {salted/in love}. embedded under a story about an amorous peanut. this finding and others by the studies cited above suggest that pragmatic information can override a low-level lexical effect. under this consideration, the negation-sensitive n400 in nieuwland & kuperberg (2008) can be said to stem from the overriding effect of pragmatic information. at first glance, the current assumption about the relationship between pragmatics and lexical semantics seems to conflict with previous findings by palaz et al. (2020). combining a pseudoword learning paradigm with a truth-value judgment task, palaz et al. (2020) conducted an erp experiment to examine the effect of the informative use of negation on an n400. in their design, 5our discussion is consistent with the original claim made by nieuwland & kuperberg (2008) that pragmatic effects can lead to incremental negation processing, but we resort to other studies in our discussion to avoide circularity. proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 27 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ negative sentences followed context sentences, as exemplified by (8). they used a pseudo-word for a subject (e.g., sloken or trante), and the context sentence enabled participants to learn what type of entity it referred to (e.g., tool or fish).6 the false negative sentence like (8b) was assumed to be pragmatically informative because it denied the participants’ belief about the referent of the pseudo-word (e.g., one assumes trante to be some kind of fish based on the context in (8b).) (8) a. context: i used a huge sloken to bang the nail in the gate. a sloken is not a fish. (true negative) b. context: i saw a trante with exotic gills in the national aquarium. a trante is not a fish. (false negative) (palaz et al. 2020: 4) despite the contextual manipulation, palaz et al. (2020) observed a negation-blind n400 effect: the true negative sentence (8a) elicited a larger n400 than their false version (8b). at first sight, this ‘pragmatics-insensitive’ n400 effect seemingly speaks against our assumption about the overriding effect of pragmatics. however, a closer look at behavioral results and a task design suggests that their manipulation of pragmatic informativeness was not sufficient enough to trigger the pragmatic effect. first, their behavioral results showed that the negative sentences took significantly longer to judge than the affirmative sentences in the presence of contexts. previous studies found that the processing asymmetry between the negative and affirmative sentences disappeared under appropriate pragmatic contexts (wason 1965, johnson-laird & tridgell 1972, glenberg et al. 1999) but not under inappropriate ones (experiment 2 in dale & duran 2011). given these previous findings, the higher processing load in the negative than affirmative sentences in palaz et al. (2020) suggests that their contextual manipulation could not trigger the strong effect of pragmatic information. second, their experiment had a methodological concern. palaz et al. (2020: 3) say “negation can become informative in contexts which express beliefs or predictions” because negation can correct the false beliefs that people have. we agree with this view, but it is doubtful that their experimental task properly tapped into this pragmatic effect of negation. in the task, for instance, participants were instructed to judge whether a trante is not a fish. in (8b) is plausible or implausible via button press. if this negative sentence was pragmatically informative due to it denying the participants’ belief about the referent of the word trante, the pragmatically correct response should be plausible but not implausible. however, we would judge this sentence as implausible if we were naive participants because it just contradicts what the context has described. in fact, palaz et al. (2020) also seem to share this judgment, since they categorized implausible as ‘accurate’ in this case. their behavioral results further suggest that this is a general intuition: they showed more than 88% ‘accuracy’ on the relevant condition. this strongly suggests that the participants in palaz et al. (2020) did not behave in a way that is consistent with how the negation should pragmatically work. for this reason, it is not clear whether their task was properly designed to examine the informative use of negation. a more appropriate task would be the one in which participants read a sentence 6palaz et al. (2020) conducted a control experiment with real words. the results were similar to those from the experiment with the pseudo-words. they can be explained under our lexical priming account. proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 28 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ and do a secondary task which does not involve plausibility judgments (e.g., a word recognition task; nieuwland & kuperberg 2008). these two considerations raise the possibility that the results of palaz et al. (2020) do not reflect the informative use of negation. therefore, the pragmatics-insensitive n400 pattern observed in palaz et al. (2020) is not problematic for our assumption about the effect of pragmatics. to conclude, removing the lexical priming relation between subject and object led to no n400 effect. although we must interpret a statistically non-significant result with caution, this suggests that the putative negation-blind n400 effect resulted from lexical priming. if this is correct, it does not necessarily constitute evidence for a strict two-step account. references bentin, shlomo, gregory mccarthy & charles c wood. 1985. event-related potentials, lexical decision and semantic priming. electroencephalography and clinical neurophysiology 60(4). 343–355. 10.1016/0013-4694(85)90008-2. carpenter, patricia a & marcel a just. 1975. sentence comprehension: a psycholinguistic processing model of verification. psychological review 82(1). 45. 10.1016/0010-0285(72)90019-9. christensen, ken ramshøj. 2020. the neurology of negation: fmri, erp, and aphasia. in viviane déprez & m teresa espinal (eds.), the oxford handbook of negation, 725–739. oxford: oxford university press. 10.1093/oxfordhb/9780198830528.013.47. clark, herbert h & william g chase. 1972. on the process of comparing sentences against pictures. cognitive psychology 3(3). 472–517. 10.1016/0010-0285(72)90019-9. dale, rick & nicholas d duran. 2011. the cognitive dynamics of negated sentence verification. cognitive science 35(5). 983–996. 10.1111/j.1551-6709.2010.01164.x. dien, joseph. 2010. the erp pca toolkit: an open source program for advanced statistical analysis of event-related potential data. journal of neuroscience methods 187(1). 138–145. 10.1016/j.jneumeth.2009.12.009. dien, joseph. 2012. applying principal components analysis to event-related potentials: a tutorial. developmental neuropsychology 37(6). 497–517. 10.1080/87565641.2012.697503. dien, joseph. 2024. multi-algorithm artifact correction (maac) procedure part one: algorithm and example. biological psychology 188. 108775. 10.1016/j.biopsycho.2024.108775. dien, joseph & gwen a frishkoff. 2005. principal components analysis of event-related potential datasets. in handy todd c (ed.), event-related potentials: a methods handbook, 189–208. mit press cambridge. dudschig, carolin, barbara kaup, mingya liu & juliane schwab. 2021. the processing of negation and polarity: an overview. journal of psycholinguistic research 50(6). 1199–1213. 10.1093/oxfordhb/9780198830528.013.47. dudschig, carolin, ian grant mackenzie, claudia maienborn, barbara kaup & hartmut leuthold. 2019. negation and the n400: investigating temporal aspects of negation integration using semantic and world-knowledge violations. language, cognition and neuroscience 34(3). 309–319. 10.1080/23273798.2018.1535127. filik, ruth & hartmut leuthold. 2008. processing local pragmatic anomalies in fictional contexts: evidence from the n400. psychophysiology 45(4). 554–558. 10.1111/j.1469proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 29 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 8986.2008.00656.x. fischler, ira, paul a bloom, donald g childers, salim e roucos & nathan w perry jr. 1983. brain potentials related to stages of sentence verification. psychophysiology 20(4). 400–409. 10.1111/j.1469-8986.1983.tb00920.x. glenberg, arthur m, david a robertson, jennifer l jansen & mina c johnson-glenberg. 1999. not propositions. cognitive systems research 1(1). 19–33. 10.1016/s1389-0417(99)00004-2. haase, viviana, maria spychalska & markus werning. 2019. investigating the comprehension of negated sentences employing world knowledge: an event-related potential study. frontiers in psychology 10. 2184. 10.3389/fpsyg.2019.02184. hagoort, peter, lea hald, marcel bastiaansen & karl magnus petersson. 2004. integration of word meaning and world knowledge in language comprehension. science 304(5669). 438–441. 10.1126/science.1095455. hald, lea a, esther g steenbeek-planting & peter hagoort. 2007. the interaction of discourse context and world knowledge in online sentence comprehension. evidence from the n400. brain research 1146. 210–218. 10.1016/j.brainres.2007.02.054. he, yifei, johanna sommer, silvia hansen-schirra & arne nagels. 2022. multivariate pattern analysis of eeg reveals nuanced impact of negation on sentence processing in the n400 and later time windows. psychophysiology e14491. 10.1111/psyp.14491. hunt iii, lamar, stephen politzer-ahles, linzi gibson, utako minai & robert fiorentino. 2013. pragmatic inferences modulate n400 during sentence comprehension: evidence from picture– sentence verification. neuroscience letters 534. 246–251. 10.1016/j.neulet.2012.11.044. jakobovits, leon a & wallace e lambert. 1962. mediated satiation in verbal transfer. journal of experimental psychology 64(4). 346. https://psycnet.apa.org/doi/10.1037/h0044630. johnson-laird, philip n & jm tridgell. 1972. when negation is easier than affirmation. quarterly journal of experimental psychology 24(1). 87–91. 10.1080/14640747208400271. kaup, barbara & carolin dudschig. 2007. the experiential view of language comprehension: how is negation represented? in schmalhofer franz & perfetti charles a. (eds.), higher level language proceesses in the brain. iinference and comprehension processes, 255–288. london: lawrence erlbaum associates publishers. kaup, barbara & carolin dudschig. 2020. understanding negation: issues in the processing of negation. in viviane déprez & m teresa espinal (eds.), the oxford handbook of negation, 635–655. oxford: oxford university press. 10.1093/oxfordhb/9780198830528.001.0001. kaup, barbara, jana lüdtke & rolf a zwaan. 2006. processing negated sentences with contradictory predicates: is a door that is not open mentally closed? journal of pragmatics 38(7). 1033–1050. 10.1016/j.pragma.2005.09.012. kaup, barbara, richard h yaxley, carol j madden, rolf a zwaan & jana lüdtke. 2007. experiential simulations of negated text information. quarterly journal of experimental psychology 60(7). 976–990. 10.1080/17470210600823512. kounios, john & phillip j holcomb. 1992. structure and process in semantic memory: evidence from event-related brain potentials and reaction times. journal of experimental psychology: general 121(4). 459. https://psycnet.apa.org/doi/10.1037/0096-3445.121.4.459. kutas, marta & kara d federmeier. 2011. thirty years and counting: finding meaning in the proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 30 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ n400 component of the event-related brain potential (erp). annual review of psychology 62. 621–647. 10.1146/annurev.psych.093008.131123. kutas, marta & steven a hillyard. 1984. brain potentials during reading reflect word expectancy and semantic association. nature 307(5947). 161–163. 10.1038/307161a0. metzner, paul, titus von der malsburg, shravan vasishth & frank rösler. 2015. brain responses to world knowledge violations: a comparison of stimulus-and fixation-triggered event-related potentials and neural oscillations. journal of cognitive neuroscience 27(5). 1017–1028. 10.1162/jocn𝑎00731. nieuwland, mante s, dale j barr, federica bartolozzi, simon busch-moreno, emily darley, david i donaldson, heather j ferguson, xiao fu, evelien heyselaar, falk huettig et al. 2019. dissociable effects of prediction and integration during language comprehension: evidence from a large-scale study using brain potentials. philosophical transactions of the royal society b 375(1791). 20180522. 10.1098/rstb.2018.0522. nieuwland, mante s, tali ditman & gina r kuperberg. 2010. on the incrementality of pragmatic processing: an erp investigation of informativeness and pragmatic abilities. journal of memory and language 63(3). 324–346. 10.1016/j.jml.2010.06.005. nieuwland, mante s & gina r kuperberg. 2008. when the truth is not too hard to handle: an event-related potential study on the pragmatics of negation. psychological science 19(12). 1213–1218. 10.1111/j.1467-9280.2008.02226.x. nieuwland, mante s & jos ja van berkum. 2006. when peanuts fall in love: n400 evidence for the power of discourse. journal of cognitive neuroscience 18(7). 1098–1111. 10.1162/jocn.2006.18.7.1098. palaz, bilge, ryan rhodes & arild hestvik. 2020. informative use of “not” is n400-blind. psychophysiology 57(12). e13676. 10.1111/psyp.13676. papeo, liuba & manuel de vega. 2020. the neurobiology of lexical and sentential negation. in viviane déprez & m teresa espinal (eds.), the oxford handbook of negation, 740–756. oxford: oxford university press. 10.1093/oxfordhb/9780198830528.013.44. rugg, michael d. 1985. the effects of semantic priming and word repetition on event-related potentials. psychophysiology 22(6). 642–647. 10.1111/j.1469-8986.1985.tb01661.x. smith, lee & raymond klein. 1990. evidence for semantic satiation: repeating a category slows subsequent semantic processing. journal of experimental psychology: learning, memory, and cognition 16(5). 852. 10.1037/0278-7393.16.5.852. tian, ye & richard breheny. 2019. negation. in cummins chris & katsos napoleon (eds.), the oxford handbook of experiemental semantics and pragmatics, 195–207. oxford: oxford university press. 10.1093/oxfordhb/9780198791768.013.29. wason, peter c. 1965. the contexts of plausible denial. journal of verbal learning and verbal behavior 4(1). 7–11. 10.1016/s0022-5371(65)80060-3. wiswede, daniel, nicolas koranyi, florian müller, oliver langner & klaus rothermund. 2013. validating the truth of propositions: behavioral and erp indicators of truth evaluation processes. social cognitive and affective neuroscience 8(6). 647–653. 10.1093/scan/nss042. proceedings of elm 3: 19-31, 2025 daiki asami, chao han, jacob burger, deanna dunlop, yue lu et al.: ‘negation-blind’ n400 effect disappears when lexical priming is controlled. 31 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ seeing vs. seeing that: children’s understanding of direct perception and inference reports e. emory davis & barbara landau* abstract. young children can reason about direct and indirect visual information, but fully mapping this understanding to linguistic forms encoding the two knowledge sources appears to come later in development. in english, perception verbs with small clause complements (“i saw something happen”) report direct perception of an event, while perception verbs with sentential complements (“i saw that something happened”) can report inferences about an event. in two experiments, we ask when 49-year-old english-speaking children have linked the conceptual distinction between direct perception and inference to different complements expressing this distinction. we find that, unlike older children or adults, 4-6-year-olds do not recognize that see with a sentential complement can report visually-based inference, even when syntactic and contextual cues make inference interpretations highly salient. these results suggest a prolonged developmental trajectory for learning how the syntax of perception verbs like see maps to their semantics. keywords. language acquisition; perception verbs; inference; semantics; syntax; complements 1. introduction the use of perception verbs like see or hear with different complement structures often corresponds to reporting distinct kinds of perceptual experiences. in english, perception verbs with small clause complements can only report direct perception of an event as it occurred, while perception verbs with embedded clause complements can be used to report an inference about an event without having directly perceived it, as demonstrated by (1) and (2) below: (1) john saw the book fall off the shelf. (2) mary saw that the book had fallen off the shelf. (1) is true only if john witnessed the event of the book falling off the shelf. while (2) can be true for mary in that same situation, it can also be true if mary simply walked into the room after the book fell and noticed it on the floor next to the shelf. this distinction relates to source monitoring, the ability to reflect on and distinguish between various sources of information and knowledge. differences in perceptual representations and knowledge arise not just from having different experiences with objects, but also from having different access to events as they unfold. an individual like john in (1), who directly witnesses an event as it occurs, has a different perceptual experience from an individual like mary in (2), who may have only seen the outcome of that event; even if john and mary end up with similar representations of an event of a book falling off a shelf, john’s representation is based on directly witnessing the event, while mary’s is based on an inference. there is evidence that even young children can distinguish between direct and indirect sources of perceptual information. ünal and papafragou (2019) used two picture-matching tasks to assess * we would like to thank research assistants arunima vijay, sydney sappenfield, and clara darcy for their work in making these experiments possible, as well as the members of the language and cognition lab at johns hopkins for their insightful and valuable feedback at all steps of the research process. authors: e. emory davis, johns hopkins university (emorydavis@jhu.edu) & barbara landau, johns hopkins university (landau@jhu.edu). proceedings of elm 1: 125-135, 2021 c©2021 e. emory davis and barbara landau published by the lsa with permission of the author(s) under a cc by license. 125 https://doi.org/10.3765/elm https://www.elm-conference.net/ turkishand english-speaking children’s ability to reason about perception and inference as sources of knowledge about events for both themselves and others. they found that 4-6-year-olds could use both direct visual evidence (a picture of a woman drinking for drink) and indirect visual evidence (a picture of footprints in the snow for walk) to reason about events and match photographs to verbs describing the events. specifically, in the latter example, children needed to infer the occurrence of a past event in order to connect a verb like walk to a picture of footprints; they were also able to ascribe that same reasoning to someone else, though children performed less well when attributing either direct or inferential knowledge to others. ünal and papafragou’s results stand in contrast to previous work that showed children under age 6 have difficulty with recognizing inference as a source of knowledge (pillow 1999, sodian & wimmer 1987, inter alia). ünal and papafragou suggest that children’s success in their tasks could be due to the fact that children were not required to produce or comprehend explicit verbal reports about visual access or mental states. while these results indicate that 4-6-year-olds may be able to reason about perception and inference as sources of knowledge, there is also evidence that at this age children have difficulty mapping this understanding to linguistic forms that encode the difference between perception and inference, especially in comprehension. the term “evidential strategies” (aikhenvald 2014) refers to the various systems languages use to mark knowledge sources, such as perception and inference. despite the differences between languages in the types of evidential strategies they employ, the acquisition process for them appears to be remarkably similar cross-linguistically. research has consistently shown that until around the age of six or seven, children’s command of these strategies is not fully adult-like (e.g. in tibetan; de villiers et al. 2009), and children’s comprehension lags behind their production (papafragou et al. 2007, ünal & papafragou 2016, ünal & papafragou 2018, winans et al. 2015). this asymmetry is notable not only for its robustness across languages and evidential typologies, but also because it is the reverse of many other productioncomprehension asymmetries in language – children often understand linguistic forms before they produce them (goldin-meadow, seligman & gelman 1976, gertner, fisher & eisengart 2006). english marks the difference between direct perception and inference in perceptual reports using several different (optional) syntactic strategies. one way is through verbal complements of perception verbs like see. as discussed above, perception verbs with small clause complements (e.g. john saw mary leave) report direct perception of an event, while perception verbs with embedded clause complements (e.g. john saw that mary left) can be used to report either direct perception or inference about an event on the basis of visual (or sometimes other) evidence. children face two potential challenges in learning that see can sometimes mean “infer.” first, see + sentential complement maps to two different readings – direct perception or inference – so children must figure out which one of these is appropriate in the context; they must also learn that other frames, like small clause complements, do not allow an inference interpretation. another aspect of this first challenge is that the direct perception reading may well be the primary or more accessible reading for the verb see. second, learning the inferential meaning of see requires children to abstract away from the visual perception component of the verb’s semantics to some extent. that is, see can report a belief one has about one thing (such as a book falling) as a result of perceiving of something else with entirely different visual properties (such as a book on the floor). previous research has shown that young children (around age 4) do in fact have difficulty learning some aspects of perception verb semantics, and tend to assign narrower meanings to perception verbs than older children or adults, treating verbs like see as only reporting an event proceedings of elm 1: 125-135, 2021 e. emory davis and barbara landau: seeing vs. seeing that: children’s understanding of direct perception and inference reports. 126 https://doi.org/10.3765/elm https://www.elm-conference.net/ involving direct perception with one’s eyes (landau & gleitman 1985, elli, bedny & landau 2021). in two experiments, we sought to determine whether and when young english-speaking children, who already produce perception verbs like see in both small clause and sentential complement frames (davis & landau 2020), have linked the conceptual distinction between direct perception and inference to the different complements expressing this distinction. 2. experiment 1. in experiment 1, we sought to determine whether 4-9-year-old children have mapped see in different syntactic frames to distinct kinds of perceptual experiences, and whether they recognize that perception verb utterances containing sentential complements can be true in different contexts than those containing small clause complements. 2.1. methods. we presented 36 children (4;0-9;01, m = 6;5) and six adults with eight different stories in which a first character directly perceives an event, while a second encounters visual evidence that could lead to an inference about the event. for example, in one of the stories, two children, lily and noah, leave a plate of cookies in the kitchen while they go out to play. lily comes back inside and catches her dog fido eating the cookies; fido runs out of the room and lily chases after him. noah then comes into the kitchen and finds an empty plate on the floor surrounded by cookie crumbs and paw prints. all stories were of similar length and complexity. narration of the stories and the sentences presented in the test phase (see below) were pre-recorded. the narrator did not use different voices for any of the characters. the narration was accompanied by visual depictions of the stories (figure 1). figure 1. images and transcript of narration from one of the stories in experiment 1. proceedings of elm 1: 125-135, 2021 e. emory davis and barbara landau: seeing vs. seeing that: children’s understanding of direct perception and inference reports. 127 https://doi.org/10.3765/elm https://www.elm-conference.net/ after each story, participants heard two sentences: a direct perception sentence reporting perception of the event as it occurred (e.g. “i saw fido eating the cookies”) and an inference sentence reporting an inference based on visual evidence (e.g. “i saw that fido had eaten the cookies”). they were then were asked to identify which of the two characters (e.g. lily or noah) said either the direct perception (4 trials) or inference sentence (4 trials). participants also heard two pairs of control sentences for each trial, checking for their ability to identify the characters (e.g., “i have red pants”) and remember details of the story (e.g., “i came into the kitchen and found an empty plate”). there were two presentations of the task; each had the stories in a different randomized order and both counterbalanced the order of the sentences for each pair of targets and controls. both target sentences were presented in every trial as we thought that children might need the contrast of the two frames within each trial to be successful in linking the inference sentence to the inferring character. that is, while i saw that fido had eaten the cookies can be truthfully said by either the direct perception character (lily) or the inferring character (noah), the inference interpretation is strengthened by contrasting the embedded clause frame with the small clause frame, since i saw fido eat the cookies can only be truthfully said by the direct perception character. responses were marked correct if participants attributed the direct perception sentence to the direct perception character (on the 4 trials where this was queried) and the inference sentence to the inferring character (on the other 4 trials), which was the expected pattern for adults. 2.2. results. participant performance was measured as the proportion of correct responses for each sentence type (direct perception or inference) across all stories. adults performed as expected and attributed the direct perception sentences to the direct perception characters and the inferences sentence to the inferring characters. adult performance was at ceiling (mean correct above 95%) for both sentence types, so detailed analysis of their responses was not conducted. both adults and children performed at ceiling for the control sentences. children’s performance was compared to chance (0.5 for each sentence type) using one-tailed single sample t-tests. responses were also analyzed with logistic mixed effects models using the lme4 package in r (bates et al. 2014). these models included sentence type (direct vs. inference), age, gender, task presentation, and trial number as fixed effects, and participant and story as random effects. if a model did not converge with all of these fixed effects, gender was removed, then task presentation and trial number. except for sentence type and age, none of these effects were found to be significant in any of the models that included them. children consistently attributed direct perception sentences to the direct perception characters (m = 0.91, t(35) = 14.407, p < 0.05), with no effect of age (β = 0.02, se = 0.02, p > 0.05). in contrast, there was an age effect for the inference sentences (β = 0.11, se = 0.03, p < 0.05): children under 7 years of age (n = 23, m = 5;5) were at chance in attributing inference sentences to the inferring character (m = 0.55; t(22) = 0.8941, p > 0.05), whereas children 7 and older (n = 13, m = 8;4) performed significantly above chance on the inference sentences (m = 0.96, t(12) = 17.725, p < 0.05) and were no different from adults (m = 0.96). we also examined joint performance on the two sentence types to determine whether individual children demonstrated a tendency to interpret both target sentences as having a direct perception meaning. each child’s responses were categorized as fitting one of four patterns: above chance for both sentence types (n = 22), which was the adult pattern; above chance for direct perception and below chance for inference (n = 10), which would indicate a direct perception interpretation of both targets; below chance for direct perception and above chance for inference (n = 2); and below chance for both (n = 2). a multinomial logistic regression analysis showed that proceedings of elm 1: 125-135, 2021 e. emory davis and barbara landau: seeing vs. seeing that: children’s understanding of direct perception and inference reports. 128 https://doi.org/10.3765/elm https://www.elm-conference.net/ age was a significant predictor of the two dominant response patterns (β = -0.97, se = 0.4, p < 0.05; figure 2), confirming that the older children understood the distinction between see and see that, while the younger children did not. figure 2. proportion of correct responses for direct perception vs. inference sentences in experiment 1. each point represents one participant; dotted lines show chance levels. points in upper right corners indicate distinct interpretations for the direct perception and inference sentences; points in upper left corners indicate a direct perception interpretation for both sentence types. 2.3. discussion. while children over 7 years old performed like adults on this task, and consistently attributed the inference sentences to the inferring characters, children under 7 attributed the inference sentences to the direct perception characters about half the time on average. the high rate of correct responses for the control questions indicates that participants of all ages understood the task and could follow the stories. participants also rarely made errors on the direct perception statements, even if they responded incorrectly for the inference statements. any difficulty that the younger participants had with the inference sentences, then, could not be due to issues with remembering the characters or the events in the stories. we believe there are two possible explanations for the younger children’s failure to consistently attribute the inference sentences to the inferring characters: children’s linguistic knowledge or pragmatic factors. first, younger children may have difficulty with the inference sentences because they have incomplete knowledge of the semantics and syntax of perception verbs. children who performed poorly on the inference sentences may have believed that see in any syntactic frame can only refer to direct perception. this account fits with previous research which has shown that younger children assign narrow meanings to perception verbs, treating see as only referring to perceiving with one’s eyes. one 6-year-old participant who chose the direct perception character for every inference sentence insisted on each trial that both target sentences were said by the direct perception character, suggesting that at least some children may have had this more limited semantic representation for see. a second possibility is that younger children’s responses were driven largely by pragmatic considerations, in particular the salience of the direct perception character for them, rather than their knowledge of the syntactic frames. given that the target sentences were statements about who saw what, children may have assumed that the experimenter wanted to know about the character who actually witnessed the event. alternatively, children may have ranked direct perception higher 4-6 yos 7-9 yos 0.25 0.50 0.75 1.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 inference d ire ct p er ce pt io n age (years) 4 5 6 7 8 9 proportion of correct responses for target sentences proceedings of elm 1: 125-135, 2021 e. emory davis and barbara landau: seeing vs. seeing that: children’s understanding of direct perception and inference reports. 129 https://doi.org/10.3765/elm https://www.elm-conference.net/ than inference in the set of possible interpretations for see. participants may have believed that the direct perception character was more likely to be the answer since that character had more direct perceptual experience with the event than the inferring character. this could explain why younger children’s responses skewed towards attributing both types of target sentences to the direct perception characters and almost never the inverse. additionally, the fact that the see + embedded clause sentence is acceptable for the direct perception character to say may have further strengthened children’s assumption that the character with direct visual experience was the better of the two choices. this contrasts with adults and older children, who did not make the direct perception interpretation of see with a sentential complement when it was contrasted with the small clause frame. these considerations may have given greater weight to the direct perception character in the younger children’s decision process, overriding consideration of the syntactic distinction and the implicature pragmatics that accompany the contrast of the two sentence types. our results suggest that it is not until around age seven that english-speaking children consistently make adult-like distinctions between the syntactic frames that see occurs with and the corresponding semantics – that is, knowing that “i saw something happen” is different from “i saw that something happened” and that such statements are appropriate in different contexts. however, since direct perception and inference were both potentially acceptable readings of see + sentential complement in this task, we cannot determine whether younger children’s difficulty is due to their understanding of the pragmatics of two readings, or because they do not have sufficient semantic and syntactic knowledge of see. the next experiment attempts to distinguish between these two possibilities. 3. experiment 2. in experiment 2, we tested whether younger children would accept see that for reporting inference in a truth-value judgment task designed to reduce some of the pragmatic complexities of experiment 1. the tvjt provides participants with the opportunity to make independent judgments about each of the target sentences; children only need to determine whether see that is acceptable in an inference scenario, rather than identify the best or most likely interpretation as they may have done in the previous forced-choice task. this new task also depicts only one individual whose perceptual access to the event varies across trials, rather than two individuals with different perceptual experiences in each trial, eliminating the possibility that participants would implicitly compare different individuals’ perceptual experiences and make linguistic judgments on the basis of who saw the event “better” (i.e. more directly). if children are aware that see + sentential complement licenses an inference reading, they should judge inference sentences as “right” in inference scenarios while judging direct perception sentences as “wrong”; if they think see (in either frame) can only report direct perception, they should reject both target sentences when the speaker has only seen evidence. 3.1. methods. participants were 14 adults and 23 children (4;05-6;10; m = 5;06). we included only children under 7 in this experiment since children 7 and older in experiment 1 performed like adults. the truth-value judgment task had two within-subjects factors: visual access to an event and sentence type. visual access had three conditions (see event, see evidence, and doesn’t see) and each was tested with queries about two target sentence types (direct perception and inference). six different events were presented per condition for a total of 18 trials, and in every trial both target sentence types were tested, plus one control sentence. the 18 target trials were presented to participants in one of three pre-determined randomized orders, with the orders themselves assigned randomly to participants. only one version of each event was used per presentation order. proceedings of elm 1: 125-135, 2021 e. emory davis and barbara landau: seeing vs. seeing that: children’s understanding of direct perception and inference reports. 130 https://doi.org/10.3765/elm https://www.elm-conference.net/ the order of the three test sentences (two targets plus one control) in each trial was randomized within each presentation of the task. participants were shown videos depicting an observer, mary, watching a simple event in which an actor causes a visible change of state to an object (e.g. the actor peels a banana). mary is prompted by the sound of a bell to put on a blindfold at different points during the videos, varying by the three visual access conditions (figure 3). in the see event condition (6 trials), mary sees the entire event. in the see evidence condition (6 trials), mary sees the object beforehand and evidence of the event afterwards (the peeled banana), but not the peeling event itself. in the doesn’t see condition (6 trials), mary sees the object before the event, but does not see the event or any evidence of it. participants themselves always saw the full event, including how much of the event mary watched. figure 3. timeline of videos in the three visual access conditions, showing when the observer, mary, put on and removed her blindfold relative to the event. after each event video, participants watched as mary made three statements in separate videos: a direct perception statement (e.g. “i saw someone peel the banana”), an inference statement (e.g. “i saw that someone peeled the banana”), and a control statement that could be true or false (e.g. “there was a banana” or “there was an orange”). participants were asked if mary was “right” or “wrong” after each sentence. the complements in the direct perception and inference sentences always matched the event, so that participants would judge the statements based on mary’s perception of the event rather than the felicity of the complement. 102 figure 4.14. timeline of videos in the three visual access conditions, showing when the observer, mary, put on and removed her blindfold relative to the event. 1 2 3 4 5 see event bell rings see evidence bell rings doesn’t see bell rings 6 7 8 9 see event bell rings see evidence bell rings doesn’t see bell rings proceedings of elm 1: 125-135, 2021 e. emory davis and barbara landau: seeing vs. seeing that: children’s understanding of direct perception and inference reports. 131 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.2. results. the expected adult-like responses to the target sentences for each visual access condition were based on adult performance in experiment 1, as well as another experiment not reported here, which showed that under some circumstances, adults accept both see and see that for direct perception events.1 participants with adult-like knowledge of the semantics of see were expected to judge direct perception sentences (“i saw…”) as “right” only when mary saw the event directly (see event trials); to judge inference sentences (“i saw that…”) as “right” when mary either saw the event (see event trials) or saw evidence of it (see evidence trials); and to judge both target sentences as wrong when mary did not see any aspect of the event (doesn’t see trials). we also expected participants to judge true control sentences as “right” and false ones as “wrong.” participants’ responses were coded as correct if they fit this pattern, and as incorrect if they did not. we expected that children who did not understand that see that can report inference would differ from this pattern in just one respect: they would judge inference sentences as “wrong” in the see evidence trials. all participants were at ceiling for the control sentences. the critical measure of performance on the target sentences was not participants’ overall accuracy for each sentence type, but the pattern of responses across both direct perception and inference sentences, particularly in the see event and see evidence conditions. we conducted a cluster analysis to identify response patterns using the mclust package in r (scrucca et al. 2016). the input to the cluster analysis was each participant’s proportion of correct responses (according to the expected pattern described above) for direct perception and inference sentences in the see event and see evidence trials only (four scores per participant), as these were the crucial conditions for assessing participants’ understanding of see. data from adult and child participants were analyzed together to more easily identify children whose responses were similar to those of adults (and vice versa). the optimal model had five clusters of equal variance. the full set of participant data (all responses to targets in all conditions) was then annotated with the cluster information, i.e. which cluster each participant belonged to as identified by the cluster analysis. the data were then analyzed with logistic mixed effects models using the lme4 package in r (bates et al. 2014) to evaluate the differences between clusters, that is, whether the clusters corresponded to significantly different ways of responding to the target sentences across the three conditions. the majority of adult responses fit into two patterns. nine of the adult participants gave responses consistent with our predicted pattern: they judged both target sentences as “right” in the see event condition, and only inference sentences as “right” in the see evidence condition (figure 4, cluster 1). three more adults gave responses that differed from the former group of adults only in one respect: they judged inference sentences as “wrong” in the see event condition significantly more often than the first group of adults (β = -4.33, se = 1.19, p < 0.05; figure 4, cluster 2), suggesting they treated the two target sentences as mutually exclusive, a pattern consistent with adult responses in experiment 1. as shown in figure 4 (top row), the participants in these two clusters overwhelmingly judged the inference sentences as “right” (m = 0.82) in the see evidence condition, demonstrating an understanding that see with an embedded clause can report inference. only one child clustered with the adults and gave responses that fit the predicted pattern for full knowledge of see complements. all other child participants showed no understanding that the different complements corresponded to a difference in meaning. their responses fit into three patterns. one third of children (n = 7), plus one adult, consistently judged see that as “wrong” in the see evidence 1 in that experiment, adults carried out a truth-value judgment task in which participants heard only one target sentence per trial; results showed that they judged both see and see that as correct for reporting direct perception. proceedings of elm 1: 125-135, 2021 e. emory davis and barbara landau: seeing vs. seeing that: children’s understanding of direct perception and inference reports. 132 https://doi.org/10.3765/elm https://www.elm-conference.net/ condition (m = 0.10; figure 4, cluster 3), which was significantly more often than adults who fit the predicted pattern (β = -4.66, se = 0.76, p < 0.05). this group otherwise made adult-like judgments of the direct perception and inference sentences, indicating that their only difficulty was in understanding that “i saw that…” could be used to report inference. somewhat unexpectedly, about half of the children (n = 12) judged all of the target sentences in every condition as “right” (figure 4, cluster 4), even in the doesn’t see condition where mary saw nothing. however, these children were at ceiling for the control sentences, so their incorrect responses were not due to an overall “right” bias or a total failure to understand the task. when asked follow up questions at the end of the experiment, the children in this group confirmed that mary was wearing her blindfold and could not see during the event in the see evidence or doesn’t see trials, so they were not confused or mistaken about her visual access in these trials. instead, many of these children said that mary’s see statement was correct because “i/we saw it” or because the described event did happen. taken together, this indicates a ‘realist’ interpretation of the target sentences – that is, these children judged mary’s see statements as “right” because the complement gave an accurate description of the event. the remaining children (n = 3) and one adult were at chance for the inference sentences in the see evidence condition (m = 0.54), suggesting uncertainty about their meaning (figure 4, cluster 5). children’s responses (across all clusters) were not predicted by age (β = 0.05, se = 0.18, p > 0.05). figure 4. mean proportions of correct responses for target sentences in experiment 2, grouped by response pattern. cluster 1: predicted adult cluster 2: alternative adult cluster 3: see that ¹ inference cluster 4: realist cluster 5: other adult=9 child=1 adult=3 child=0 see event see evidencedoesn't see see event see evidencedoesn't see0.00 0.25 0.50 0.75 1.00 visual access condition m ea n pr op or ti on c or re ct p er c lu st er direct perception inference adult=9 child=1 adult=3 child=0 see event see evidencedoesn't see see event see evidencedoesn't see0.00 0.25 0.50 0.75 1.00 visual access condition m ea n pr op or ti on c or re ct p er c lu st er direct perception inference target sentence direct perception “i saw...” inference “i saw that...” a du lt pa tt er ns c hi ld p at te rn s correct responses by cluster adult=9 child=1 adult=3 child=0 see event see evidence doesn't see see event see evidence doesn't see 0.00 0.25 0.50 0.75 1.00 visual access condition m ea n pr op or tio n c or re ct p er c lu st er direct perception inference adult=1 child=7 adult=0 child=12 adult=1 child=3 see event see evidence doesn't see see event see evidence doesn't see see event see evidence doesn't see 0.00 0.25 0.50 0.75 1.00 visual access condition m ea n pr op or tio n c or re ct p er c lu st er direct perception inference proceedings of elm 1: 125-135, 2021 e. emory davis and barbara landau: seeing vs. seeing that: children’s understanding of direct perception and inference reports. 133 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.3. discussion. the results of experiment 2 show that children under 7 years old do not demonstrate an awareness of the inference meaning of see + sentential complement despite linguistic and pragmatic conditions optimized to support this reading. that is, even when factors that could have biased children toward a direct perception interpretation of the target sentences were removed, most children did not show an understanding of the distinction between see with small clause and sentential complements. the majority of adults accepted see that sentences for reporting inference, but only one child did; the majority of children gave responses reflecting nonadult like interpretations of see and its complements. about a third of the children said see with a sentential complement was “wrong” in the see evidence condition; this is consistent with the results of experiment 1, in which children showed a tendency to interpret see that as having only a direct perception meaning, and suggests that linguistic knowledge, and not solely pragmatics, can account for children’s performance in that task. additionally, about half of the children in experiment 2 judged all see statements as “right” regardless of what the speaker had actually seen. in fact, 13 children commented during the task that the two target sentences were the same. furthermore, age did not predict children’s response patterns, so it was not the case that, for example, only the youngest participants in this experiment were the ‘realists.’ given that the children 7 and older in experiment 1 performed like adults, the overall lack of an age effect for children ages 4-6 suggests that there is a qualitative change in children’s understanding of see that around the age of seven. 4. conclusion. our experiments show that 4-6-year-olds do not recognize that see can report visually-based inference when it takes a sentential complement (e.g. “i saw that someone peeled the banana”), even when pragmatic and contextual cues make inference interpretations highly salient. our results indicate that adults and children over seven have both the direct perception and inference readings as part of the semantics of see, and make use of a variety of cues to select from these possible interpretations. in particular, adults and older children can use syntactic information (complement type) and pragmatic information (such as the implicature of the contrast of multiple frames and the perceptual experience shown in context) to determine when see means “perceive directly with the eyes” vs. “infer from (visual perception of) evidence.” children between four and seven, however, are still learning the syntax and semantics of perception verbs like see and how distinct syntactic forms encode different kinds of perceptual experience. our results suggest a significant change in children’s semantic representations around age seven, with earlier representations corresponding to see as encoding only direct visual perception, and later ones coming to include knowledge that see can report inference and an understanding of the relationship between a wider range of frames and their meanings. one caveat is that we only tested children’s comprehension of statements describing another individual’s experience, and judgments about what another person would infer from visual evidence could be particularly difficult. children might show an understanding of inferential see when asked about their own inferences rather than someone else’s, a possibility we are examining now. even so, our results are consistent with cross-linguistic findings on children’s acquisition of evidential language, which has also shown a protracted developmental trajectory, particularly in the comprehension of forms used to report others’ perception or knowledge. previous work has suggested that this difficulty in learning markers of direct perception vs. inference, regardless of marking type or language being acquired, may be the result of children’s difficulty integrating reasoning about other people’s experiences with their own linguistic knowledge (ünal & papafragou 2018, de villiers et al. 2009). thus, the potentially universal qualitative shift in proceedings of elm 1: 125-135, 2021 e. emory davis and barbara landau: seeing vs. seeing that: children’s understanding of direct perception and inference reports. 134 https://doi.org/10.3765/elm https://www.elm-conference.net/ children’s understanding of such language around age seven may reflect a change in their cognitive ability to synthesize linguistic and conceptual information about perception and knowledge. references aikhenvald, alexandra y. 2004. evidentiality. oxford university press. bates, douglas, martin mächler, ben bolker & steve walker. 2014. fitting linear mixedeffects models using lme4. journal of statistical software 67(1). 1–48. davis, e. emory & barbara landau. 2020. seeing and believing: the relationship between perception and mental verbs in acquisition. language learning and development. 1–21. https://doi.org/10.1080/15475441.2020.1862660. de villiers, jill, jay garfield, harper gernet-girard, tom roeper & margaret speas. 2009. evidentials in tibetan: acquisition, semantics, and cognitive development. (ed.) s. a. fitneva & t. matsui. evidentiality: a window into language and cognitive development, new directions for child and adolescent development 125. 29–47. https://doi.org/10.1002/cd.248. elli, giulia, marina bedny & barbara landau. 2021. how does a blind person see? developmental change in applying visual verbs to agents with disabilities. cognition 212. 104683. https://doi.org/10.31234/osf.io/dpgx5. gertner, yael, cynthia fisher & julie eisengart. 2006. learning words and rules. psychological science 17(8). 684–692. https://doi.org/10.1111/j.1467-9280.2006.01767.x. goldin-meadow, susan, martin e.p. seligman & rochel gelman. 1976. language in the twoyear old. cognition 4(2). 189–202. https://doi.org/10.1016/0010-0277(76)90004-4. landau, barbara & lila r. gleitman. 1985. language and experience: evidence from the blind child. cambridge, ma: harvard university press. papafragou, anna, peggy li, youngon choi & chung hye han. 2007. evidentiality in language and cognition. cognition 103. 253–299. https://doi.org/10.1016/j.cognition.2006.04.001. pillow, bradford h. 1999. children’s understanding of inferential knowledge. the journal of genetic psychology 160(4). 419–428. https://doi.org/10.1080/00221329909595555. scrucca, luca, michael fop, t. brendan murphy & adrian e. raftery. 2016. mclust 5: clustering, classification and density estimation using gaussian finite mixture models. r journal 8(1). 289–317. https://doi.org/10.32614/rj-2016-021. sodian, beate & heinz wimmer. 1987. children’s understanding of inference as a source of knowledge. child development 58(2). 424. https://doi.org/10.2307/1130519. ünal, ercenur & anna papafragou. 2016. production-comprehension asymmetries and the acquisition of evidential morphology. journal of memory and language 89. 179–199. https://doi.org/10.1016/j.jml.2015.12.001. ünal, ercenur & anna papafragou. 2018. relations between language and cognition: evidentiality and sources of knowledge. topics in cognitive science. 1–21. https://doi.org/10.1111/tops.12355. ünal, ercenur & anna papafragou. 2019. how children identify events from visual experience. language learning and development. psychology press 15(2). 138–156. https://doi.org/10.1080/15475441.2018.1544075. winans, lauren, nina hyams, jessica rett & laura kalin. 2015. children’s comprehension of syntactically-encoded evidentiality. proceedings of nels 45. 189–202. proceedings of elm 1: 125-135, 2021 e. emory davis and barbara landau: seeing vs. seeing that: children’s understanding of direct perception and inference reports. 135 https://doi.org/10.3765/elm https://www.elm-conference.net/ corpus evidence for the role of world knowledge in ambiguity reduction: using high positive expectations to inform quantifier scope noa attali, lisa s. pearl, & gregory scontras* abstract. every-negation utterances (e.g., every vote doesn’t count) are ambiguous between a surface scope interpretation (e.g., no vote counts) and an inverse scope interpretation (e.g., not all votes count). investigations into the interpretation of these utterances have found variation: child and adult interpretations diverge (e.g., musolino 1999) and adult interpretations of specific constructions show considerable disagreement (carden 1973, heringer 1970, attali et al. 2021). can we concretely identify factors to explain some of this variation and predict tendencies in individual interpretations? here we show that a type of expectation about the world (which we call a high positive expectation), which can surface in the linguistic contexts of every-negation utterances, predicts experimental preferences for the inverse scope interpretation of different every-negation utterances. these findings suggest that (1) world knowledge, as set up in a linguistic context, helps to effectively reduce the ambiguity of potentiallyambiguous utterances for listeners, and (2) given that high positive expectations are a kind of affirmative context, negation use is felicitous in affirmative contexts (e.g., wason 1961). keywords. scope ambiguity; universal quantifiers; negation; pragmatics; computational models; corpus linguistics; psycholinguistics 1. introduction. it’s unclear how people prefer to interpret ambiguous every-negation utterances, such as (1): (1) every vote doesn’t count. a. no vote counts. surface scope: ∀x[vote(x) →¬count(x)] (every > n’t) b. not all the votes count. inverse scope: ¬∀x[vote(x) → count(x)] (n’t > every) the utterance in (1) allows both a surface interpretation (1a) and an inverse one (1b), depending on the logical scope of the quantifier relative to negation. which interpretation would a listener or reader believe is more likely to be intended by the speaker? previous studies show variation in interpretation behavior. on the one hand, children and non-native speakers seem to disprefer the inverse scope interpretation of every-negation (musolino 1999, gualmini et al. 2008, viau et al. 2010, chung & shin 2022); converging research on scope ambiguity suggests that surface scope is easier to access, involving less representational complexity or processing cost (tunstall 1998, pritchett & whitman 1995, anderson 2004, lee et al. 2011). on the other hand, experimental studies that directly measure interpretation preference by adult native english speakers find a preference for inverse scope interpretations of everyand all-negation (carden 1970, heringer 1970, carden 1973). *we would very much like to thank the quantitative language collective at uc irvine for helpful comments and feedback on this work. authors: noa attali, uc irvine (nattali@uci.edu), lisa s. pearl, uc irvine (lpearl@uci.edu) & gregory scontras, uc irvine (g.scontras@uci.edu). proceedings of elm 2: 13-23, 2023 c©2023 noa attali, lisa s. pearl and gregory scontras published by the lsa with permission of the author(s) under a cc by license. 13 https://doi.org/10.3765/elm https://www.elm-conference.net/ additionally, the above studies report preferences that are aggregated across different sentences and contexts. when we look beyond average interpretations to behavior on individual sentences, we find further variation. changes in the immediate linguistic context – in fact, in the sentence containing the quantifier-negation construction itself – can flip interpretation patterns. carden (1973) found that for (2), 82.5% of respondents said that only the inverse scope interpretation was possible, 7.5% said that both were possible but that they favored inverse scope, and none reported accessing only surface scope. on the other hand, for (3), 100% said that only the surface scope interpretation was possible. (2) all the boys didn’t arrive, did they? (3) all the boys didn’t leave until midnight. this variation highlights how findings on interpretation patterns depend on the stimuli themselves. what then are the characteristics of the contexts that impact these interpretation preferences for every-negation utterances? that is, when are particular scope interpretations preferred? in their computational cognitive model of scope ambiguity, scontras & pearl (2021) demonstrate that a kind of world knowledge we term a “high positive expectation” (hpe) can explain some variation in interpretation preferences with every-negation utterances. to illustrate, an hpe for every vote doesn’t count is the prior belief that it’s highly likely that every vote does count. that is, the hpe for this every-negation utterance is that the worlds consistent with the equivalent positive utterance (every vote does count) have a high probability. in general, an hpe is the prior belief that the most likely true world states are those consistent with the every world state in the universe of the utterance. an hpe could contribute to the felicity of using every-negation with an inverse scope not all interpretation (e.g., not all the votes count), thereby reducing the ambiguity of the utterance for listeners. this ambiguity reduction occurs because the not all interpretation is highly informative about the expectation that all is true (by indicating that this expectation isn’t correct). we might put it this way: in a context with this kind of expectation (e.g., all votes count), the inverse scope interpretation of every-negation (e.g., not all votes count) is felicitous as an emphatic message that the salient expectation (e.g., that all votes do in fact count) is false. as scontras & pearl (2021) suggest, this world knowledge factor is one way to quantitatively specify a pragmatic factor that explains prior behavioral results (see scontras & pearl 2021 for more details). scontras & pearl’s model implements a cooperative, efficient speaker – cooperative in wanting to say something true, and efficient in wanting to be informative. for example, returning to the counting votes case, the model predicts that when the context holds the expectation that all votes count, speakers would more likely agree that every vote doesn’t count could be used to mean not all votes count. given this hpe context, the inverse scope interpretation of every vote doesn’t count felicitously – i.e., cooperatively and efficiently – conveys that some votes do not, in fact, count. in general, the model predicts that speakers tend to endorse every-negation (e.g., every vote doesn’t count) as a true description of a scenario consistent with the inverse scope interpretation (e.g., not all votes count) when every-negation conveys that an hpe (e.g., all votes count) is false (e.g., it’s false that all votes count). the model’s predictions for the role of an hpe extend well to prior findings. attali et al. (2021) proceedings of elm 2: 13-23, 2023 noa attali, lisa s. pearl and gregory scontras: corpus evidence for the role of world knowledge in ambiguity reduction. 14 https://doi.org/10.3765/elm https://www.elm-conference.net/ find that scontras & pearl’s model successfully predicts listener interpretations of every-negation when given an hpe. specifically, the modeled listener shows a close qualitative and quantitative match to the average preference for the inverse scope interpretation of an out-of-context sentence every marble isn’t red. the inverse scope not all interpretation is preferred over surface scope none when given an hpe because (1) there are more ways for the not all interpretation to be true compared with the none interpretation, and so a cooperative speaker intends that interpretation, and (2) it’s highly informative to update a strongly biased, salient belief such as an hpe (e.g., that every vote does count), and so an efficient speaker intends that interpretation; for more details see attali et al. 2021). because hpes can explain interpretations for every-negation utterances in prior experimental and computational work, we ask how well hpes account for interpretations of every-negation utterances in naturalistic contexts. for instance, do hpes surface in the contexts of spontaneously produced every-negation? specifically, when a local linguistic context seems to express an hpe, is an inverse scope interpretation more likely than a surface scope interpretation? we first describe how we developed a corpus of naturally-occurring every-negation uses in context via a behavioral study, including the preferred interpretation of each use. we then describe how we identified hpes in the corpus contexts. we present our analysis for the connection between the presence of an hpe and an item’s preferred interpretation, finding that inverse scope interpretations are indeed more preferred following an hpe. our results suggest that the world knowledge implemented as an hpe in the local linguistic context can help effectively reduce the ambiguity of every-negation utterances for listeners, and that negation use is felicitous in affirmative contexts, like the kind that hpes encode. 2. corpus data and behavioral experiment. to assess the role of hpes in the interpretation of naturally-occurring every-negation utterances, we need (1) a corpus of naturally-occurring everynegation utterances in context, (2) a measure of these utterances’ preferred interpretations, and (3) an estimate of the extent to which a context contains an hpe. in this section, we describe how we created a corpus (goal 1) and measured interpretation preferences in context (goal 2). section 3 discusses how we identified hpes in context (goal 3) to see whether their presence predicts inverse scope interpretations. 2.1. corpus search for every-negation utterances. to achieve goal 1, we identified uses of every-negation in a corpus of spontaneous speech. we extracted the every-negation occurrences in the speech section of the corpus of contemporary american english (coca; davies 2015), which is made up of transcripts of spoken conversations from american radio and tv programs from 1990 to 2012 (≈9 million clauses, or ≈95 million words). we defined every-negation occurrences as those where quantified subjects precede and c-command sentential negation (with not or contracted n’t). to develop the automated search, we randomly selected a year of coca transcripts and manually searched it for uses of every-negation. we then wrote a search that yielded a 100% recall rate, that is, returning each of the occurrences in this development set. we applied this search to the rest of the coca speech section. we found that every-negation uses are highly infrequent in english conversation; in total, we identified 390 cases. proceedings of elm 2: 13-23, 2023 noa attali, lisa s. pearl and gregory scontras: corpus evidence for the role of world knowledge in ambiguity reduction. 15 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.2. corpus annotation. to achieve goal 2, following degen (2015), we crowd-sourced interpretation preferences of these uses in their immediate contexts (three preceding sentences and one following sentence). for each item, participants (n = 208) completed a paraphraseendorsement task (scontras & goodman 2017), choosing on a sliding scale between none and not all paraphrases of the potentially-ambiguous clause. 2.2.1. participants. we recruited 390 participants with u.s. ip addresses through amazon.com’s mechanical turk (mturk) crowd-sourcing service. of these 390, we assessed data from 208 participants (35% female; mean age: 41 y/o) who passed control trials (discussed further in section 2.2.4) and indicated that english was their only native language. each received $2.00. 2.2.2. stimuli. for each of the 390 every-negation uses, we created excerpts consisting of the three preceding context sentences, the bolded potentially-ambiguous clause, and one following context sentence (see figure 1). we also created paraphrases of the surface and inverse scope interpretations. as part of a pilot experiment (n=94), we checked that the paraphrase wording was correctly understood, so that the surface scope interpretation paraphrase was always understood to be consistent with a none situation (e.g., that none are red describes three blue marbles rather than two red and one blue marble) and the inverse scope paraphrase was always understood to be consistent with a some but not all situation (e.g., that not all are red preferentially describes two red and one blue marble rather than three blue marbles). because the ambiguous clauses took the form quantified noun phrase–verb–negation–remainder (e.g., everybody is not doing that: quantified noun phrase = everybody; verb = is; negation = not; remainder = doing that), surface scope paraphrases took the form none/no one/nobody/nothing– verb–remainder (e.g., nobody is doing that) and inverse scope paraphrases took the form not all/not all things are–remainder (e.g., not all are doing that). figure 1: sample paraphrase-endorsement trial from the corpus analysis of every-negation. 2.2.3. design. the initial instructions asked to “choose the best paraphrase for the bolded part” for fifteen randomly-selected conversation excerpts; at each trial, participants were again asked “what did the speaker mean in the bolded part?” (see figure 1). beneath the excerpt, participants rated the best paraphrase as a judgment on a sliding scale between the surface and inverse scope interpretations. the two scope interpretations were randomly assigned for each item in left-right or right-left order. proceedings of elm 2: 13-23, 2023 noa attali, lisa s. pearl and gregory scontras: corpus evidence for the role of world knowledge in ambiguity reduction. 16 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.2.4. controls. to check that participants were reading and understanding the contexts of the items – and also as a way to suggest that context is useful – two control trials were constructed to imitate the items from the corpus. these control trials contained clearly disambiguating information about the intended scope interpretation in the surrounding context. the disambiguating information always appeared as a restatement of the speaker’s meaning. these two controls appeared in random order as the first two trials for each participant. the surface scope-disambiguating control item is in (4), and the inverse scope-disambiguating control item is in (5). for clarity, the disambiguating information is italicized, though it was not italicized in the experiment. participants were considered to pass the surface control by placing the slider closer to the none paraphrase than to the not all paraphrase; they passed the inverse control by placing the slider closer to the not all paraphrase than to the nobody paraphrase. (4) tonhauser: the ten board members voted last night. i was really surprised—i thought at least some of them would like proposition 23. but all ten of them voted against it. basically, every board member didn’t like proposition 23. not even a single one of them liked it. (5) sidner: look, we completely fixed the issue. indicators have improved across the board. everybody’s happy. grosz: (voiceover) no, everybody isn’t happy. some are happy but others are deeply dissatisfied with what they call a ‘band aid solution.’ we restricted analysis to those participants who passed both controls and indicated that english was their only native language. the rate of passing both controls was 53%. this relatively high failure rate may have been due to low english reading proficiency. though we restricted mturk participation to us ip addresses and to those mturk workers who have completed at least 1,000 tasks in the past and we also only analyzed data from self-reported native english speakers, it’s possible that some participants didn’t fluently read english well enough. another factor may have been attention and motivation: participants in an online platform may be disengaged with the experiment. a third factor is task difficulty: perhaps the paraphrase endorsement task could be seen as a complex reading comprehension and logical inference task, because these sentences have multiple logical operators. 2.3. results. each item was judged by at least 2 and at most 14 different participants. although the surface scope paraphrases randomly appeared on the left or right of the sliders, we transformed and report responses on sliders as though the surface scope paraphrases always appeared on the left. this allows the final response measure to vary from 0 (maximum endorsement of the surface scope interpretation) to 1 (maximum endorsement of the inverse scope interpretation). in the coca transcripts, we found both a general preference for inverse scope interpretations as well as a high degree of interpretation variation for the every-negation utterances. figure 2, which shows judgment-by-judgment interpretations, suggests that many of these utterances in context seem unambiguous: 29% of individual scores were below 0.25 (indicating a strongly surface scope interpretation) while 53% of individual scores were above 0.75 (indicating a strongly inverse scope interpretation). figure 3 aggregates ratings by items. for some items, strong intuitions are reliable across different participants’ judgments: 12% of mean scores were below 0.25, and 38% of mean scores were above 0.75. examples (6) and (7) are items that elicited a strong proceedings of elm 2: 13-23, 2023 noa attali, lisa s. pearl and gregory scontras: corpus evidence for the role of world knowledge in ambiguity reduction. 17 https://doi.org/10.3765/elm https://www.elm-conference.net/ surface scope preference ((6): mean response ≈ 0.13) or a strong inverse scope preference ((7): mean response ≈ 0.98). figure 2: individual scope interpretations from the every-negation corpus analysis. figure 3: mean interpretations per item from the every-negation corpus analysis. (6) @!wertheimer: so what about new jersey? can new jersey get over the hump? @!rapoport: well, [transcript cuts out] first team from the eastern division to return to the finals since the bulls were winning all their championships. they’re a little nervous about that in new jersey, linda, that every team that made it to the finals from the east in the last couple of years hasn’t been able to repeat; but again, they’re strong competition. the pistons have been impressive this year. a. no team (that made it to the finals from the east in the last couple of years) has been able to repeat. (every > n’t) b. not all teams (that made it to the finals from the east in the last couple of years) have been able to repeat. (n’t > every) (7) @!caller hi. my question for mr. eisner was, mgm is one of my favorite places in disneyworld and one of my favorite attractions there is the animation studios, and now the studio, the animation studio there is closed, and everything has moved to california, and i proceedings of elm 2: 13-23, 2023 noa attali, lisa s. pearl and gregory scontras: corpus evidence for the role of world knowledge in ambiguity reduction. 18 https://doi.org/10.3765/elm https://www.elm-conference.net/ wanted to know how you justified doing that. @!eisner well, everything has not moved to california. we will still be demonstrating animation in florida. a. nothing has moved to california. surface scope (every > not) b. not all things have moved to california. inverse scope (not > every) 2.4. discussion. the results of the paraphrase endorsement study with the corpus-mined stimuli show variation and a weak inverse scope preference in adult native english speakers’ interpretations for every-negation utterances. although these results agree with the larger picture painted by previous studies on every-negation, to our knowledge this is the first larger-scale investigation of naturalistic stimuli in context. in the following section, we ask whether hpes account for this variation and inverse scope interpretation preference. more specifically, when an inverse scope interpretation is preferred for an every-negation utterance (e.g., not all votes count for every vote didn’t count), was that use of every-negation in fact an emphatic message that a salient hpe (e.g., all votes count) is false? 3. identifying high positive expectations in linguistic contexts. an hpe represents a prior belief, and one way to measure for its presence is by its overt linguistic expression in an item’s preceding context. for example, for every vote doesn’t count, an hpe is the prior belief that every vote does count – that is, that the worlds consistent with the non-negated utterance (every vote does count) are highly probable – and one measure of this hpe’s presence is the non-negated utterance itself: every vote does count. as a preliminary measure, the first author hand-coded categorically for the presence/absence of an overt hpe expression in each preceding context. we found that 59/390 (15%) of the items contained such an expression. for an automatic and more objective measure of the hpe expression – that is, a method that could scale to large amounts of data and would capture the intended linguistic phenomenon while minimizing experimenter bias – we calculated the degree of lexical overlap between the preceding linguistic context and a string representing the positive expectation (pos exp). for each item (e.g., every vote doesn’t count), we first coded pos exp as the potentially-ambiguous clause without negation (e.g., every vote does count). we then coded for the extent to which the pos exp appeared in the preceding context as the longest common substring similarity (lcs; needleman & wunsch 1970) between each preceding context string c and pos exp pair, calculated using the r stringdist package (loo 2014). each lcs was equal to the longest sequence formed by pairing words from the preceding context string c and pos exp, while keeping their order intact; the dissimilarity dlcs(c, pos exp) was then the number of unpaired words left over in both strings. dlcs(c, pos exp) can be defined recursively as in (8) for different relative lengths of the two strings to be matched against: (i) it is trivially 0 for empty strings (line 1: c = pos exp = ϵ). (ii) it is based on pairing each word from both strings if the two strings have equal length (line 2: |c| = |pos exp|). for example, suppose the preceding context is every vote does count for an utterance with the pos exp every vote does count. then, dlcs= 0. however, if the preceding context was what is going on?, dlcs= -8 because all eight words in the two strings would be unpaired. (iii) it is based on the minimum lcs-distance that can be obtained from pairing all the words proceedings of elm 2: 13-23, 2023 noa attali, lisa s. pearl and gregory scontras: corpus evidence for the role of world knowledge in ambiguity reduction. 19 https://doi.org/10.3765/elm https://www.elm-conference.net/ from the shorter string to an equal number of words from the longer string (line 3: otherwise). for example, suppose the preceding context was i believe every vote does count for an utterance with the pos exp every vote does count. then, dlcs= -2 because all four words in pos exp would pair to every vote does count in the context, and leave unpaired the two words i believe. thus, dissimilarity ranges from 0 (completely similar) to the total words w in both strings combined (completely dissimilar), where w = (|c| + |pos exp|). we calculate lcs similarity as negative dissimilarity: −dlcs(c, pos exp). thus, lcs similarity ranges from 0 to -w , with values closer to zero indicating more lexical overlap. in particular, values closer to zero indicate a greater similarity between the context and the hpe linguistic string, and so represent a higher probability that the context contained a linguistic string transparently encoding an hpe. (8) dlcs(c, pos exp) =    0, if c = pos exp = ε dlcs(c1:|c|−1, pos exp1:|pos exp|−1), if |c| = |pos exp| 1 +min{dlcs(c1:|c|−1, pos exp), dlcs(c, pos exp1:|pos exp|−1)}, otherwise. 4. results. 4.1. hand-coded hpe results. using the preliminary categorical hand-coding where we found that 59/390 of the utterances had hpes, we first looked at p(inverse|hpe): how often an inverse scope interpretation was preferred when an hpe occurred. we found that 50/59 (85%) of utterances with hpes were on average better paraphrased by the inverse scope paraphrase not all than the surface scope paraphrase none. we also looked at p(hpe|inverse) vs. p(hpe|surface): how often items where the inverse interpretation was strongly preferred had an hpe compared with items where the surface interpretation was strongly preferred. we found that 22% of highly inverse-preferred items (those with responses greater than 0.75) had hpes, as opposed to 6% of highly surface scope-preferred items (those with responses less than 0.25). these results suggest that the hand-coded hpes do tend to co-occur with an inverse scope interpretation in our sample. however, the automatic measure of an hpe’s presence that we described above allows us to to measure the continuous relationship between the extent of hpe expression and the strength of inverse scope preference, as shown below. in addition, this measure can be used in future work to analyze larger samples. 4.2. automatic hpe results. we used the continuous lcs-based measure −dlcs to assess if an hpe predicts an inverse scope preference per item, and ran a linear mixed effects model predicting logit-transformed mean item responses by −dlcs (representing lcs similarity) with random intercepts for participants (see figure 4). to determine whether an hpe captures individual judgment variation above and beyond mean item-level variation, we predicted logit-transformed item responses by lcs similarity, with random intercepts for participants and items. both models found that lcs similarity was a significant predictor of an inverse scope preference preference (p < .001 in both). that being said, although lcs similarity is a significant predictor of inverse scope preference, the relationship is noisy, as figure 4 shows, with a marginal r2 = 0.024. proceedings of elm 2: 13-23, 2023 noa attali, lisa s. pearl and gregory scontras: corpus evidence for the role of world knowledge in ambiguity reduction. 20 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 4: preceding hpe and average inverse scope item preference. interestingly, only preceding, and not following, expressions of hpes predict an inverse scope preference, as figure 5 shows: a version of both models that calculated lcs similarity using overlap with the following – rather than preceding – context, found lcs similarity of the following context not to be a significant predictor of either item-level or judgment-level interpretations. figure 5: following hpe and average inverse scope item preference. 5. discussion. our corpus analysis suggests that a high positive expectation (hpe) expressed directly in the preceding linguistic context can affect scope interpretation preferences for everynegation utterances. in particular, hpes expressed this way correlate with stronger preferences for the inverse interpretation. these results align with previous modeling results (scontras & pearl 2021) and pragmatically-oriented proposals from truth-value judgment studies for supporting the proceedings of elm 2: 13-23, 2023 noa attali, lisa s. pearl and gregory scontras: corpus evidence for the role of world knowledge in ambiguity reduction. 21 https://doi.org/10.3765/elm https://www.elm-conference.net/ felicity of every-negation in context (gualmini et al. 2008). in particular, an hpe provides a context that makes an inverse scope interpretation more felicitous because several things hold: (1) the listener assumes that speakers are truthful and informative, (2) an hpe represents a strong belief about the world, (3) finding out an hpe is false (that not all is true) is very informative, and (4) the inverse scope not all interpretation is one way to find this out. we speculate that perhaps in comparison with alternative constructions such as not every (e.g., not every vote counts for every vote doesn’t count), the every-negation construction highlights that a positive expectation is false, and might even be preferred as more informative (in such a context) than its not every alternative. we note that our lcs similarity measure for hpes is a first-pass one (the first anyone has tried to our knowledge), and likely underestimates hpe presence. in particular, this measure looks for transparent linguistic encodings of an hpe; but of course world expectations do not have to be encoded linguistically, or encoded nearby even if they are linguistically encoded. even given our restriction to overtly expressed world knowledge in the preceding three sentences, our specific measurement of lcs similarity has a noisy potential to underestimate the presence of an hpe for several reasons. first, it is affected by context length, such that lcs similarity is lower for longer contexts even if those contexts contain a clear hpe expression. for instance, returning to the vote-counting case, lcs similarity would be -2 for the context i believe every vote does count but it would be 0 the context every vote does count. second, this lcs similarity measure looks for an hpe based on the exact lexical items in the every-negation utterance. for instance, it would identify the hpe in the context every vote does count for the every-negation utterance every vote doesn’t count; yet, this measure misses the hpe in the context all votes should matter because the individual lexical items differ (every vs. all, count vs. matter). this rigidity of lcs similarity as a measure of context-sentence overlap could be a source of the noisiness evidenced in figure 4. other potential sources of noise include additional factors that may help disambiguate scope interpretations, such as prior expectations about questions under discussion and grammatical scope accessibility (e.g., as found by scontras & pearl 2021). still, the advantage of lcs similarity is that it provides an automatic continuous measure to improve our analysis of larger-scale data. here, it allowed us to consider the potential linear relationship between the extent of hpe expression and the extent of an inverse scope preference. future work could replace lcs similarity with a measure that considers a vectorized semantic representation of meaning rather than lexical overlap between the context and a string representing the hpe. a vectorized semantic measure would allow for the flexibility to recognize degrees of semantic similarity rather than categorical lexical equivalence. for example, such an approach would allow us to count all votes should matter as a context expressing an hpe for every vote doesn’t count (recognizing that all is similar to every, and count similar to matter, in this context). more generally, our results suggest that listeners can rely on world knowledge and properties of the immediately surrounding contexts of an ambiguous utterance, like an every-negation utterance, to interpret it. since a context containing an hpe is a kind of affirmative context, these findings support the broader theory that negation use is more felicitous in affirmative contexts (e.g., wason 1961). one way that listeners understand every-negation in context is as a kind of emphatic frame for the message that an hpe is false. proceedings of elm 2: 13-23, 2023 noa attali, lisa s. pearl and gregory scontras: corpus evidence for the role of world knowledge in ambiguity reduction. 22 https://doi.org/10.3765/elm https://www.elm-conference.net/ references anderson, catherine. 2004. the structure and real-time comprehension of quantifier scope ambiguity: northwestern university evanston, il dissertation. attali, noa, gregory scontras & lisa s pearl. 2021. pragmatic factors can explain variation in interpretation preferences for quantifier-negation utterances: a computational approach. proceedings of the annual meeting of the cognitive science society 43(43). carden, guy. 1970. a note on conflicting idiolects. linguistic inquiry 1(3). 281–290. carden, guy. 1973. disambiguation, favored readings, and variable rules. new ways of analyzing variation in english 171–82. chung, eun seon & jeong-ah shin. 2022. native and second language processing of quantifier scope ambiguity. second language research 1–26. davies, mark. 2015. corpus of contemporary american english (coca) https://doi.org/ 10.7910/dvn/amuduw. degen, judith. 2015. investigating the distribution of some (but not all) implicatures using corpora and web-based methods. semantics and pragmatics 8. 11–1. gualmini, andrea, sarah hulsey, valentine hacquard & danny fox. 2008. the question–answer requirement for scope assignment. natural language semantics 16(3). 205. heringer, james t. 1970. research on quantifier-negative idiolects. chicago linguistic society 6. 95. lee, miseon, hye-young kwak, sunyoung lee & william o’grady. 2011. processing, pragmatics, and scope in korean and english. japanese/korean linguistics 19. 297–311. loo, mark p.j. van der. 2014. the stringdist package for approximate string matching. the r journal 6(1). 111–122. 10.32614/rj-2014-011. https://doi.org/10.32614/ rj-2014-011. musolino, julien. 1999. universal grammar and the acquisition of semantic knowledge: an experimental investigation into the acquisition of quantifier-negation interaction in english: university of maryland, college park dissertation. needleman, saul b & christian d wunsch. 1970. a general method applicable to the search for similarities in the amino acid sequence of two proteins. journal of molecular biology 48(3). 443–453. pritchett, bradley & john whitman. 1995. syntactic representation and interpretive preference. japanese sentence processing 65–76. scontras, gregory & noah d goodman. 2017. resolving uncertainty in plural predication. cognition 168. 294–311. scontras, gregory & lisa s pearl. 2021. when pragmatics matters more for truth-value judgments: an investigation of quantifier scope ambiguity. glossa: a journal of general linguistics 6(1). tunstall, susanne lynn. 1998. the interpretation of quantifiers: semantics & processing: university of massachusetts at amherst dissertation. viau, joshua, jeffrey lidz & julien musolino. 2010. priming of abstract logical representations in 4-year-olds. language acquisition 17(1-2). 26–50. wason, peter c. 1961. response to affirmative and negative binary statements. british journal of psychology 52(2). 133–142. proceedings of elm 2: 13-23, 2023 noa attali, lisa s. pearl and gregory scontras: corpus evidence for the role of world knowledge in ambiguity reduction. 23 https://doi.org/10.3765/elm https://www.elm-conference.net/ de re interpretation in belief reports—an experimental investigation yuhan zhang & kathryn davidson* abstract. determiner phrases (dps) under intensional operators (e.g., want, believe, must, may) give rise to multiple interpretations, known as the de re/de dicto ambiguity. formal theoretical approaches to modeling this ambiguity must rely on nuanced semantic judgments, but inconsistent judgments in the literature suggest that informal judgment collection may be insufficient. in addition, little is known about how these ambiguities are resolved in a context and how preferences between these readings vary by contexts and across individuals, etc. we reported three controlled experiments to systemize the truth-value judgment collection of de re/de dicto readings. while the de dicto readings were robustly accepted by nearly all english speakers, de re readings exhibited strongly bimodal judgments, suggesting an inherent disagreement among speakers. in addition, the acceptability of de re judgments was affected by the dp’s internal structure as well as idiosyncratic scenarios. more broadly, our experimental results lend support to the practice of including quantitative data collection within semantics. keywords. de re/de dicto ambiguity; truth-value judgment; quantitative method 1. introduction. in natural language semantics, the de re/de dicto ambiguity captures the referential properties of determiner phrases (dps) in intensional domains1. when the intensional operator is an attitude predicate (e.g., want, believe), dps read de re refer via a description that holds in the context of the speaker; the attitude holder introduced by the target sentence needs not commit to the referential association between the linguistic expression and its referent. in contrast, dps read de dicto have descriptions that need only hold in the possible worlds introduced by the intensional operator associated with the attitude holder, and consequently, the referential association might not hold for the speaker. for a concrete example, first consider the intensional predicate want and the indefinite dp a prince: in the sentence aurora wants to marry a prince with the de re reading, the speaker communicates that there is a specific prince that aurora wants to marry, even if aurora does not recognize him as a prince, while under the de dicto reading, the speaker would be informing us that aurora wants to marry whoever is a prince. (the first is true in the world of sleeping beauty; the second is false.) a (mostly, although not entirely) parallel situation happens with definite dps: the sentence aurora wants to marry the prince, under what we will call its de re reading, is true when aurora’s desires include marrying a spe * many thanks to shannon bryant, gennaro chierchia, judith degen, masoud jasbi, joshua martin, giuseppe riccardi, jack robinovitch, uli sauerland, jesse snedeker, julia sturm, and audiences at the m&m linguistics laboratory, harvard langcog workshop, and elm 1 for critical feedback (mistakes are our own). thanks to yilan wang for the assistance with the experiment pictures. data collection was supported by a graduate student research grant from institute for quantitative social science at harvard university awarded to yz. authors: yuhan zhang (yuz551@g.harvard.edu) and kathryn davidson (kathryndavidson@fas.harvard.edu), harvard university. 1 as early as aristotle, linguistic phenomena related to de re/de dicto have been observed. yet this pair of latin terms was not intensively applied until the medieval period by thomas aquinas. the adoption of the terms in philosophy and linguistics was initiated by frege, russell, and quine but the current sense of de re and de dicto is not directly or intuitively related to the literal latin meaning of the terminology (de re: ‘of the thing’, de dicto: ‘of what is said’) (von fintel & heim, 2011). therefore, it is clearer to introduce the de re/de dicto distinction via contextualized examples. for more details of the nomenclature, see keshet and schwarz (2019). proceedings of elm 1: 310-321, 2021 c©2021 yuhan zhang and kathryn davidson published by the lsa with permission of the author(s) under a cc by license. 310 https://doi.org/10.3765/elm https://www.elm-conference.net/ cific person who is a prince in the actual world, even if she does not realize that he is the prince. what we will call here its de dicto reading is true when aurora’s desires include “marrying the prince”, even if she is mistaken about who the prince may be. in this paper, we examine current semantic judgment collection methods for formal theories of the de re/de dicto ambiguity. we report three novel experiments which highlight some advantages of obtaining quantitative judgments for readings with this ambiguity, focusing on definite dps, given that they have already generated some judgment inconsistency in the literature. in the first section, we lay out the existing theoretical landscape and raise the motivation for the quantitative approach. in the second section, we report three controlled truth-value judgment experiments that tested whether the inconsistency also existed among multiple native english speakers. in the final section, we endeavor to account for the observed judgment variation. 1.1. a brief overview of current theoretical frameworks. a traditional approach to the de re/de dicto distinction models it as a scope ambiguity: in the logical form de re (but not de dicto) dps outscope the intensional operator to obtain the reading assigned in the actual world, as in (1) (see russell 1905, fodor 1970, cresswell & von stechow 1982 among others). (1) aurora wants to marry a prince. de re lf: ∃x[princew0(x) ù "w[belw0(aurora, w) ⟶ [marryw(aurora)(x)]]] de dicto lf: "w[belw0(aurora, w) ⟶ ∃x[princew(x) ù marryw(aurora)(x)]] while this theory can capture the distinction for simple (indefinite) dps as in (1), when they appear under quantification the outscoped de re reading diverges nontrivially from the intended reading. in response, percus (2000) proposes a solution using situation pronouns. every verb phrase (vp) and noun phrase (np) takes a situation pronoun as an unpronounced intensional variable that is bound by a higher lambda abstractor to get its world assignment. in this way, dps can attain the de re interpretation via the logical binding of intensional variables while remaining in-situ. however, situation pronouns overgenerate, predicting readings not attested in natural language, resulting in a proposal by keshet (2008, 2011) known as split intensionality, a return to the traditional scope-based theory with the employment of a type-shifting operator associated with an intensional operator. raising a dp above this operator but below the intensional operator not only makes the dp an intensional argument, but also assigns it a different world index from the matrix clause. consequently, dps can be interpreted de re once they are raised above the type-shifting operator while remaining in the scope of the intensional operators. however, neither the scope-based theories nor the situation pronoun theory can adequately account for the “multiple-guise” scenario raised by quine (1956), and consequently the theory of de re readings has been enriched by the addition of concept generators (see (4) below), permitting de re readings without movement (percus & sauerland 2003, anand 2006, charlow & sharvit 2014). the de re/de dicto distinction illustrates how formal theories evolve in response to new pieces of linguistic observation. importantly, these linguistic observations not only include the well-formedness of a sentence, but also its truth-value judgment offered by linguists given a corresponding scenario. given the complexity of the data (especially the context/scenario) involved in de re/de dicto judgements, we find it unsurprising that there is some judgment inconsistency from existing publications. we review some inconsistencies in the next section, and conclude that they call for a more consistent judgment collection approach, an approach that generates, in tonhauser and matthewson (2015)’s term, “stable, replicable, and transparent” judgments for observations of the same kind in the collection of such layered semantic data. proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 311 https://doi.org/10.3765/elm https://www.elm-conference.net/ 1.2. the need for quantitative research. inconsistencies found in linguistics and relevant fields constitute our motivation for the current study. in the linguistic theory literature, there exist some inconsistent de re judgments about the same dp structure in nearly identical linguistic environments. for instance, von fintel and heim (2011) employ sentence (2) to argue that the dp your abstract exhibits genuine ambiguity of de re and de dicto interpretations, while nelson (2019) argues that the target dp her brother in (3), which has the same internal structure as your abstract, cannot be interpreted de re. (2) john believes that [your abstract]de re will be accepted. de re truth condition: john reviewed an amazing abstract and thought that it will be accepted. the speaker of this belief report has the additional knowledge that the abstract is written by the addressee “you” and thus utters (2), while john does not know the authorship of the abstract. (von fintel and heim 2001:157) (3) # sally believes that [her brother]de re is happy. (supposed) de re truth condition: sally hears a person laughing outside on the street who happens to be her brother. she believes that the person is happy, even though she does not recognize him as her brother. (nelson 2019:13) nelson’s reason to deny the de re reading is that the belief holder sally does not conceive of the person—sally’s brother in real world—as her brother. nelson takes sally’s perspective and argues that de re should not be true given the scenario, while von fintel and heim claim that de re readings are a regular natural language phenomenon. one may wonder whether nelson’s reasoning is representative of a dispreference for de re: to test this, it helps to gather judgment data across a wider variety of scenarios. or, perhaps particular linguistic features, e.g., your vs. her, or the inanimate vs. animate possessees, contribute to the unexpected judgment inconsistency. or, perhaps individuals vary in their preferred readings when faced with ambiguity. controlled scenarios, examples, and a broader population allow us to test these hypotheses. further inconsistency in the linguistics literature arises in more complex cases. charlow and sharvit (2014) note one such disagreement: while they claim that the possessee mother in (4) should be de dicto, they report that another linguist who works on the same phenomenon finds it more natural under a de re reading. (4) john believes that [every female studenti]de re likes [heri]de re [mother]de dicto. lf: john believes-w0 [λ8 λ9 λ1 [every female student-w0 [λ2[[g8 t2]-w1 likes-w1 [g9 her2]-w1 mother-w1]]]] truth condition of a “multiple-guise” scenario: john comes into contact with every actual female student more than once, and each actual female student appears each time in a different guise. the same woman appears in two different guises and john fails to recognize this. he thinks he came into contact with two different women. furthermore, in john’s mind, the mapping between the different guises is one-to-one. the sentence is about a specific scenario when john believes that a likes b’s mother, c likes d’s mother, and e likes f’s mother when in reality, a = b, c = d, and e = f. proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 312 https://doi.org/10.3765/elm https://www.elm-conference.net/ outside of theoretical linguistics, real-life instances provide data that are occasionally unexpected given theoretical predictions. while it is claimed that cardinal dps cannot be interpreted de re (musan 1995, keshet 2008, romoli & sudo 2009), sentence (5), an utterance collected at an economic conference reported on language log by liberman (2005), suggests otherwise. (5) u.s. forces in iraq have intentionally killed [12 journalists]de re. de re truth condition: there were 12 journalists killed by the u.s. forces in an attack but the forces did not know the people they killed were journalists. (liberman on october 23, 2005) moreover, outside linguistics, researchers in related fields (i.e., philosophy, psychology, law) have claimed that the de re reading is easier to obtain in scenarios where both de re and de dicto are admitted, which is not (as far as we know) a claim that has been made within linguistics. in philosophy, jaszczolt (1997) maintains that the directly referential property of definite noun phrases is more salient in communication and thus argues for a “default de re reading”. her perspective finds its allies in cognitive science and developmental psychology. for example, when a participant and a character/protagonist in an experiment both know the identity of an object but the protagonist remains partially ignorant of the object’s certain properties, both children and adult participants fail to restrict their description to the properties already known by the protagonist to refer to the object when put into the protagonist’s shoes (mitchell et al. 1996, apperly & robinson 2003, apperly & butterfill 2009, low & watts 2013). these observations suggest an egocentrism or reality bias explanation and the easiness of accessing information in actual reality but not others’ mental status may bias one to expect something like a “default de re” hypothesis. this bias is also bolstered in legal settings where the focus on a literal interpretation of the defendant’s action rather than his intention to conduct such action during jury procedure echoes the “default de re” claim (anderson 2013). despite the observed judgment inconsistency in linguistics literature and the “default de re” claim outside linguistics, there has been no experimental work, that we are aware of, that directly looks at the judgment preferences for de re/de dicto readings in a given scenario. while hackl et al. (2009) has studied transparent versus opaque readings in intensional transitive predicates using online reading times and gathered evidence supporting the scope-based theory over the situation pronoun approach, their finding—qred transparent dps facilitate the processing of the following acd site—does not directly address the judgment inconsistencies introduced above. fortunately, crowdsourcing techniques offer a systematic solution. by gathering offline judgments from multiple native speakers, we can detect whether the observed inconsistency results from idiosyncratic noises or inherent disagreement; by creating multiple scenarios to test a single phenomenon, we can confirm whether the judgment is robust to more variation. crucially, by comparing the de re and de dicto readings of the same sentence under minimally different contexts, we can attempt a fair comparison of the judgment pattern and hopefully understand factors involved in different preferences for these readings. in sum, there is good reason to believe that the de re/de dicto literature can benefit from more quantitative methods. 1.3. research goal. we aim to set up a simple and efficient experimental template for systematically obtaining stable, replicable, and transparent judgments of de re/de dicto readings across carefully controlled scenarios. we focus on definite dps since several above examples with questionable judgments are definite (e.g., your abstract, her mother), although we note that this definite non-de re reading differs from the traditional de dicto exhibited by indefinite dps. proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 313 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2. experiment one. experiment 1 used highly controlled scenarios that permitted both de re and de dicto readings of definite dps to probe its judgment pattern from native english speakers. 2.1. participants. 120 adult native english speakers were recruited through amazon’s mechanical turk. they received $2 compensation for finishing the experiment. 2.2. materials and design. in this experiment, participants read four written scenarios and gave their judgment of a target sentence based on each scenario (an example in table 12). context julie is one of the judges of an ongoing poetry competition. the best poem that she has read so far is an extremely intriguing poem about the ocean. she believes that this poem will win the competition. julie remembers being told that nicole, one of the bestknown poets, submitted a poem about the ocean to the competition. therefore, julie concludes that this poem must be written by nicole and the first prize will be going to her. however, this poem was actually written by elizabeth, a younger and lesser-known poet. it is just a coincidence that the two poets wrote about the same topic. judgment question according to this story, please use the slider bar to indicate to what extent you agree or disagree with the following statement. target sentence i: julie believes that elizabeth’s poem will win the competition. (de re) target sentence ii: julie believes that nicole’s poem will win the competition. (de dicto) table 1: example scenario in exp.1 in each scenario, there were two terms that described the target object (e.g., poe m). the protagonist (e.g., julie) associated one term x (e.g., nicole’s poem) with the target object but in reality, x was not correct and the correct descriptive term y (e.g., elizabeth’s poem) was not known by the protagonist. if the wrong term x was used in reporting the protagonist’s belief, a de dicto reading would emerge; if the correct term y was used a de re reading would emerge. given this scenario, both readings were predicted to be true (e.g., romoli & sudo 2009). in each scenario, the participants read one of the two target sentences (target sentence i or target sentence ii, varied between participants) and dragged a slider bar to show the extent to which they agree or disagree. after the participants’ decision, a numeric judgment score was recorded (from “highly agree” = 100 to “highly disagree” = −100). three sanity check sentences were additionally provided in each scenario—one was definitely correct, one was definitely wrong, and the last was uncertain. successful judgments on these sentences indicated the partici 2 the full experiment material and the statistical analysis in exp.1-3 can be accessed via https://osf.io/qgnr5/. the online survey can be accessed via https://harvard.az1.qualtrics.com/jfe/form/sv_01utaqo9hkkacgf. proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 314 https://doi.org/10.3765/elm https://www.elm-conference.net/ pants were attentive and thus eligible for inclusion in the data analysis. the advantage of a slider bar is its greater sensitivity to reveal potential judgments that would otherwise stay concealed due to the strong categorical implication in designs like the binary or the likert scale (marty et al. 2020). each participant read four scenarios and each scenario was coupled with four sentences for judgment elicitation. two of the four scenarios were randomly chosen for the de re condition and the other two for the de dicto condition. the participants were randomly assigned to one of six lists created for latin square design. the order of the four stories was randomized, as was the order of the four sentences within each scenario. the entire survey was created on qualtrics and distributed on amazon’s mechanical turk. 2.3. results. we analyzed only the responses from participants who correctly judged the correct and incorrect sanity checks at least 75% of the time. 115 participants’ data (95.8%) were retained for the analysis. the histogram in figure 1 shows the judgment distribution of de re/de dicto sentences across all scenarios. while judgments for de dicto readings overwhelmingly aggregate toward the “highly agree” end, judgments for de re readings are bimodal—although more than half of the judgments are agreed with, another sizable proportion goes to the “highly disagree” edge. we further analyzed the agreement proportion in each scenario, assuming it was appropriate to treat the continuous judgment as a binary variable given its categorial distribution. in figure 2, the de dicto agreement rates are at ceiling for all scenarios while de re judgments exhibit larger variability across scenarios with a unanimously lowering effect (𝜒2 = 79.13, df = 1, p < .001). figure 1 & 2: histogram distribution of judgments; agreement rate on condition and scenario3 the visual difference of de re/de dicto condition was confirmed by a mixed-effects logistic regression analysis. by treating the de re/de dicto conditions and the scenario as sum-encoded fixed effects with a random intercept on participants, we found that the de re trials were significantly less likely to be agreed with compared with a random trial (β = −1.61, se = 0.23, p < .001); additionally, scenario b had a significantly higher agreement rate (β = 0.96, se = 0.30, p = .001) while scenario c had a significantly lower one (β = −0.81, se = 0.24, p < .001)4. 3 the error bars in figure 2 represent 95% confidence intervals sampled via random bootstrapping. 4 the syntax of the model is “logit(agree) ~ condition + context + (1|subject)”. we also ran another model “logit(agree) ~ condition * context + (1|subject)” but didn’t find significant improvement from the reported one. the models were run using r with package lme4. the reason to treat scenario as a fixed effect rather than a random effect was that it was also of theoretical interests and that its number (n = 4) was not eligible for a random effect. proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 315 https://doi.org/10.3765/elm https://www.elm-conference.net/ additionally, we explored whether groups of participants had different judgment behavior, exhibited in table 2. clearly, a preponderance of participants agreed with both de dicto trials while the judgment behavior for de re had three representative groups, suggesting that an inherent disagreement among speakers or an unplanned scenario effect may both contribute to this distinctive participant behavior of interpreting de re.5 2.4. discussion. by setting up scenarios that theoretically allow both de re and de dicto readings of definite dps and eliciting native speakers’ judgments, we found that while de dicto readings were unanimously available to participants, de re readings exhibited bimodal judgments with larger variations across scenarios and speakers. the sizable disagreement proportion and bimodal pattern suggest systematicity within the previously observed inconsistency in the literature. # of participants agree with 0 trial agree with 1 trial agree with 2 trials total de re 21 (18.3%) 45 (36.5%) 52 (45.2%) 115 de dicto 0 (0.0%) 7 (6.1%) 108 (93.9%) 115 table 2: proportion of participants by the judgment behavior 3. experiments two and three. while experiment 1 probed the judgement distribution for de re and de dicto readings of definite dps (in particular, possessive constructions) in relatively simple sentences, experiments 2 and 3 asked whether the judgment disparity could extend to other dp structures or more sophisticated sentences. driven by such kind of motivation, we studied the nuanced case of bound de re observed in charlow and sharvit (2014) for sentences like john believes that every female studenti likes heri mother (above in (4)). the crucial bound de re assigns the qnp every female student and the possessive pronoun her a de re reading. the reading of mother is less critical for theoretical choices, but given these sentences were reported to have inconsistent judgments, we decided it was also worth investigating. 3.1. participants. 160 participants in experiment 2 and 128 in experiment 3 took the tasks for $2 on amazon’s mechanical turk. after applying the same filter as in experiment 1, 127 participants (79.38%) in experiment 2 and 120 (93.75%) in experiment 3 contributed to the analysis6. 3.2. methods. we treated the de re/de dicto reading of qnp as a between-experiments manipulation (de re in experiment 2 and de dicto in experiment 3) so that the possessive pronoun and the possessee, serving as two within-subjects manipulations, took either de re or de dicto readings within one experiment. focusing on a 2 x 2 within-subjects manipulation within an experiment prevented participants from reading more than four complex scenarios and getting fatigued. the general design was nearly the same as experiment 1. the most significant difference is that while experiment 1 had each scenario allow both de re and de dicto readings, experiments 2 and 3 were designed such that each scenario supported (i.e., made true) one reading, with the target sentence held constant and the scenario manipulated across conditions. there were also illustrative pictures to facilitate processing (see an example in table 3). furthermore, experiments 2 and 3 had the same randomization, counterbalance, and filler design as experiment 1. 5 thanks to alexander göbel for pointing to this inter-speaker investigation. 6 the lower inclusion rate in experiment 2 was because those participants were recruited on weekends when it is more challenging to gather good data via online implementation. proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 316 https://doi.org/10.3765/elm https://www.elm-conference.net/ context as a photographer, john likes to rearrange his collections of photographs. one day, he encounters two sets of photos. in the first set, there are three ladies and each is holding a baby. john naturally believes that each of the babies is being held by their mother. in the second set, three young adults are each wearing a t-shirt with a “2018” logo. john naturally believes that they were graduating students in the year 2018. john also notices that, interestingly, each of the young adults shares a similar smile to one of the ladies in the first photo set. he tries to recall if there is a connection between the young adults and the ladies but memory fails him. as a matter of fact, what john fails to recall is three pieces of information. (1) the young adults in the second set of photos were the babies in the first set. they’ve grown up! (2) the ladies in the first set are actually the babies’ grandmother. the three young adults inherit their smile from their grandma who is mistakenly believed by john to be their mother. (3) the second set of photos were taken not in the graduation ceremony but when the three adults were volunteering for an academic conference in 2018. despite the fact that john doesn’t remember the correct relationship between the ladies in the first set of photos and the young adults in the second set and that he has incorrect information, john spends some time appreciating these photos. judgment question according to this story, please use the slider bar to indicate to what extent you agree or disagree with the following statement. target sentence: looking at his photos, john believes that [every conference volunteer]de re in the second set has the same smile as [their]de re [mother]de dicto. table 3: the canonical bound de re scenario, adapted from charlow & sharvit (2014) 3.3. results. figure 3 shows that consistent with experiment 1, judgments tend to gather around both scale ends. crucially, there is still a salient proportion of disagreement for the bound de re case compared with the “control” condition where all three nominal constructions were de dicto. figure 4 shows the agreement rate of the eight conditions in eight columns. the second proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 317 https://doi.org/10.3765/elm https://www.elm-conference.net/ column represents the canonical bound de re condition whose agreement rate is slightly above the chance level. overall, the de re condition of the possessee leads to a lower agreement rate than the de dicto condition (𝜒2 = 23.63, df = 1, p < .001). figure 5 displays the agreement rates by the condition manipulation and the scenarios. a visual examination shows a clear scenario effect because of the conspicuous lowering agreement rate in scenario 4 whose peculiarity is neither expected nor designed. figure 2: agreement distribution in experiments 2 and 3 figure 4 & 3: agreement proportion among conditions (4); and scenarios (5) the effects of de re/de dicto manipulation were further analyzed via a logistic mixed-effects model. the maximal model had one random intercept on participants and three fixed-effects variables to indicate the de re/de dicto assignment of the three nominal terms. the fourth fixedeffect variable was the story plus an interaction term between the story and each of the three nominal terms7. all the fixed effects were sum-coded. the results show that while de re qnps did not significantly affect the agreement rate (β = −0.06, se = 0.10, p = .561), de re possessive pronouns (β = −0.19, se = 0.08, p = .02) and de re possessees (β = −0.42, se = 0.08, p < .001) did significantly lower the rate. compared with the 7 the syntax of the model is “logit(agree) ~ (qnp + pronoun + possessee)*story + (1|subject)”. we’ve also fitted another model “logit(agree) ~ qnp + pronoun + possessee + story + (1|subject)” but the anova() test showed the former one outperformed the latter. bound de re control proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 318 https://doi.org/10.3765/elm https://www.elm-conference.net/ average agreement rate across the four scenarios and controlling the reading assignments of the three terms, scenario 2 (β = 0.96, se = 0.15, p < .001) and 3 (β = 0.41, se = 0.14, p = .003) were more likely to be agreed upon (more analysis on https://osf.io/qgnr5/). 3.4. discussion. experiments 2 and 3 tested the judgment of de re and de dicto in the bound de re type sentence and found that the canonical bound de re structure ([qnp]de re, [possessive pronoun]de re, [possessee]de dicto) was agreed with more than 50% of the time, but at the same time (and, like de re readings generally) obtained salient disagreements. in general, experiments 2 and 3 replicated the finding of experiment 1 in showing that de re readings lead to bimodal judgment. an additional result here is that the agreement rate of de re appears dependent on the dp’s internal structure and/or position—the effect of de re/de dicto variation on the agreement rate was significant for the possessive pronouns and possessees but nearly negligible for qnps. furthermore, the judgment was also affected by specific scenarios (e.g., peculiar scenario 4). furthermore, the salient difference in de re and de dicto readings of the possessee doesn’t support the challenge that a de re possessee is more natural, but rather is in line with charlow and sharvit (2014)’s main argument. the lack of effect for qnps echoes numerous observations in theoretical work that both de re and de dicto readings for qnps are felicitous (e.g., mary 1978, keshet 2008, romoli & sudo 2009 among others). going back to the claim of bound de re in charlow and sharvit (2014), these two experiments suggest that (a) bound de re does exist for many speakers, but also that (b) not everyone agrees. 4. conclusion and discussion. in a series of three experiments, we provided some quantitative evidence that there is a truth-value judgment disparity between de re and de dicto readings of definite dps, both in simple attitude reports and more complex ones. in particular, we showed that there is a systematic inconsistency for de re judgments: the bimodal distribution is far from the uniform or normal distribution, and we note that this pattern of bimodal agreement would have been impossible to detect without quantitative methods that used response options more sensitive than a binary true/false (here, we used a slider bar). this inconsistency occurred not only in experiment 1 where the scenarios admitted both readings, but also in experiments 2 and 3 where the complex scenarios were controlled to admit only a single interpretation: even when it was the only one supported/true in the given scenario, de re readings had bimodal acceptance while de dicto readings were overwhelmingly accepted. other relevant factors that also appear to affect judgments include the internal structure of the dps (e.g., possessive, quantificational, etc.) and features of the idiosyncratic scenarios. while de re readings of combinations of other types of dp structures (including, crucially, indefinite dps) and other kinds of attitude reports and intensional operators await testing, it is worth considering what may have led to disagreements in experiments 1 to 3. one cause might be that participants possess different grammars or dialects and one variation disallows the de re reading. this has to do with grammar variation and without information about the participants’ linguistic profile, this claim stays as a speculation. another cause might be related to the scenario setup where the juxtaposition of de re and de dicto terms in the written scenario enhances participants’ sensitivity to which term is used for reference in which possible world. the contrastive information evaluated in two parallel worlds could be well tracked by the participants and thus when their incremental comprehension starts from, for example, julie believes that…, there is a chance that they only attend to what julie believes and subsequently to descriptive terms held true in julie’s belief world. a de dicto dp naturally matches what the belief holder expects and thus is highly agreed upon, while encountering a de re dp whose referential relation to the entity is not held in julie’s mind could raise disagreement (as the case in example (3) raised by nelson proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 319 https://doi.org/10.3765/elm https://www.elm-conference.net/ (2019)). this sensitivity to the contrast also alludes to theory of mind ability and reasoning with perspective shifting (wimmer & perner 1983, apperly & butterfill 2009, low & watts 2013). for further investigation along these lines, it may be informative to experimentally test whether the contrast of de re and de dicto terms in the scenario influences de re judgment. that is, if there is no contrastive term such as nicole’s poem vs. elizabeth’s poem in one scenario but just one de re term unknown to the belief holder, what might be the de re agreement rate? if agreement increases, then we may conclude that the contrastive information may contribute to the high disagreement proportion of de re. to conclude, our findings highlight the value of including quantitative methods as the basis for theoretical work, especially when the linguistic observation in question raises inconsistent judgments, as in the case of de re readings for definite dps. the essential advantage of quantitative research is that with multiple speakers, multiple scenarios, and controlled manipulation, it is possible to detect whether an inconsistency observed from limited cases (e.g., sentences (2) and (3)) is noise or true disagreement, and whether a preference for one reading is due to a minor contrast/preference or a grammatical unavailability. here, we have uncovered evidence that inconsistencies about the de re readings were not due to noise or uncertainty among participants (which would lead to more intermediate agreement responses) but rather to systematic bimodal judgments and affected by scenario and dp-specific factors. we speculate that our findings may be especially enlightening in the case of known semantic/scope ambiguities. what, then, to do with such results is an important question for theoreticians. we end by noting that these results did not even require researchers to create a massive number of scenarios to uncover these patterns: in experiments 1 to 3, a mere four scenarios—just “a little bit experimental” in davidson (2020)’s term—were enough to observe this systematic disagreement, which held across each of the experiments. lastly, we hope that this work will lead not just to more work along the de re/de dicto line, but contribute to the growing field of experiments in linguistic meaning (elm). references anand, pranav. 2006. de de se. phd thesis, massachusetts institute of technology. anderson, jill c. 2013. misreading like a lawyer: cognitive bias in statutory interpretation. harvard law review 127(6). 1–74. apperly, ian a. & butterfill, stephen a. 2009. do humans have two systems to track beliefs and belief-like states? psychological review 116(4). 953–970. https://doi.org/10.1037/a0016923 apperly, ian a. & robinson, e. j. 2003. when can children handle referential opacity? evidence for systematic variation in 5and 6-year-old children’s reasoning about beliefs and belief reports. journal of experimental child psychology 85(4). 297–311. https://doi.org/10.1016/s0022-0965(03)00099-7 charlow, simon & sharvit, yael. 2014. bound “de re” pronouns and the lfs of attitude reports. semantics and pragmatics 7(3). 1-43. https://doi.org/10.3765/sp.7.3 cresswell, maxwell j. & von stechow, arnim. 1982. de re belief generalized. linguistics and philosophy 5(4). 503–535. https://doi.org/10.1007/bf00355585 davidson, kathryn. 2020. is “experimental” a gradable predicate? proceedings of north east linguistics society (nels) 50. fodor, janet d. 1970. the linguistic description of opaque contexts. phd thesis, massachusetts institute of technology. hackle, martin, koster-moeller, jorie, & gottstein, andrea. 2009. processing opacity. proceedings of sinn und bedeutung 13. 171–185. proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 320 https://doi.org/10.3765/elm https://www.elm-conference.net/ jaszczolt, kasia m. 1997. the ‘default de re’ principle for the interpretation of belief utterances. journal of pragmatics 28(3). 315–336. https://doi.org/10.1016/s0378-2166(97)00006-4 keshet, ezra. 2008. good intensions: paving two roads to a theory of the de re / de dicto distinction. phd thesis, massachusetts institute of technology. keshet, ezra. 2011. split intensionality: a new scope theory of de re and de dicto. linguistics and philosophy 33(4). 251–283. https://doi.org/10.1007/s10988-011-9081-x keshet, ezra & schwarz, florian. 2019. de re/de dicto. in jeanette gundel & barbara abbott (eds.), the oxford handbook of reference, 167–202. oxford university press. https://doi.org/10.1093/oxfordhb/9780199687305.013.10 liberman, mark. 2005. rarely better than de re. blog. language log. http://itre.cis.upenn.edu/~myl/languagelog/archives/002573.html low, jason & watts, joseph. 2013. attributing false beliefs about object identity reveals a signature blind spot in humans’ efficient mind-reading system. psychological science 24(3). 305– 311. https://doi.org/10.1177/0956797612451469 marty, paul, chemla, emmanuel, & sprouse, jon. 2020. the effect of three basic task features on the sensitivity of acceptability judgment tasks. glossa: a journal of general linguistics 5(1). 72.1–23. https://doi.org/10.5334/gjgl.980 may, robert c. 1978. the grammar of quantification. phd thesis, massachusetts institute of technology. mitchell, p., robinson, e. j., isaacs, j. e., & nye, r. m. 1996. contamination in reasoning about false belief: an instance of realist bias in adults but not children. cognition 59(1). 1–21. https://doi.org/10.1016/0010-0277(95)00683-4 musan, renate. 1995. on the temporal interpretation of noun phrases. phd thesis, massachusetts institute of technology. nelson, michael. 2019. the de re/de dicto distinction (supplement to propositional attitude reports). in edward n. zalta (ed.), the stanford encyclopedia of philosophy. metaphysics research lab, stanford university. https://plato.stanford.edu/archives/spr2019/entries/propattitude-reports/dere.html percus, orin. 2000. constraints on some other variables in syntax. natural language semantics 8(3). 173–229. https://doi.org/10.1023/a:1011298526791 percus, orin & sauerland, uli. 2003. on the lfs of attitude reports. proceedings of sinn und bedeutung 7. 228–242. quine, willard v. 1956. quantifiers and propositional attitudes. the journal of philosophy 53(5). 177–187. https://doi.org/10.2307/2022451 romoli, jacopo & sudo, yasutada. 2009. de de/de dicto ambiguity and presupposition projection. proceedings of sinn und bedeutung, 13. russell, bertrand. 1905. on denoting. mind 14(56). 479–493. tonhauser, judith & matthewson, lisa. 2015. empirical evidence in research on meaning. manuscript. von fintel, kai & heim, irene. 2011. intensional semantics. manuscript. http://lingphil.mit.edu/papers/heim/fintel-heim-intensional.pdf wimmer, heinz & perner, josef. 1983. beliefs about beliefs: representation and constraining function of wrong beliefs in young children’s understanding of deception. cognition 13. 103–128. proceedings of elm 1: 310-321, 2021 yuhan zhang and kathryn davidson: de re interpretation in belief reports—an experimental investigation. 321 https://doi.org/10.3765/elm https://www.elm-conference.net/ are second language speakers more pragmatically tolerant? explaining the differences in scalar implicature generation between l2 and l1 irene mognon, amber l. marree & petra hendriks* abstract. children’s difficulties with scalar implicature (si) generation have been argued to stem from their tolerance towards pragmatic violations rather than from issues with the inferential process per se (katsos & bishop 2011). ternary judgment tasks have been used to support this view. in these tasks, when presented with underinformative sentences, children, as well as adults, choose an intermediate option between acceptance and rejection, thus demonstrating sensitivity to underinformativeness. some recent studies show that adult second language (l2) speakers also generate sis at lower rates. in this work, we investigated whether pragmatic tolerance, possibly emerging because of limited language exposure, could explain the difference between (adult) l2 and l1 speakers. contrary to our expectations, neither our l1 control group nor our l2 groups (l2 high and l2 low proficiency) consistently selected the intermediate option when judging underinformative sentences. however, the l2 low proficiency group showed a significantly higher tendency to accept underinformative sentences compared to the l1 group. hence, our results do not support the hypothesis that l2 speakers are more pragmatically tolerant than l1 speakers. however, our findings show that, despite the adoption of a ternary judgment task, low-proficient l2 speakers display a strong tendency to interpret underinformative sentences literally. we argue that this tendency in the l2 can be attributed to the increased cognitive effort involved in si generation. keywords. scalar implicatures; pragmatics; pragmatic tolerance; l2 1. introduction. natural language utterances can often receive more than one interpretation. consider, for instance, a sentence with the quantifier some like the one presented in (1): (1) some of my plants need water. given its syntactic and semantic features, the literal interpretation of (1) corresponds to (2): (2) at least some and maybe all of my plants need water. that (2) is the literal interpretation of (1) can be verified by observing that (1) does not appear to be incompatible with a context in which all the plants need water. despite this, however, (1) can also be interpreted as in (3): (3) not all of my plants need water. unlike (2), the interpretation in (3) is argued to arise via scalar implicature (si) generation (grice 1989, geurts 2010, noveck 2018). according to the traditional approach, sis are inferences based on the linguistic alternatives that the speaker could, but did not use in producing the utterance. specifically, si generation can be described as follows: in order to communicate that all the plants need water, sentence (4), rather than (1), should be used. * authors: irene mognon, university of groningen and goethe university frankfurt (mognon@em.unifrankfurt.de), amber l. marree, university of groningen (a.l.marree@student.rug.nl) & petra hendriks, university of groningen (p.hendriks@rug.nl). proceedings of elm 3: 236-249, 2025 c©2025 irene mognon, amber l. marree, and petra hendriks published by the lsa with permission of the author(s) under a cc by license. 236 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (4) all of my plants need water. thus, when the speaker utters (1) and not (4), (1) can be taken to imply that, according to the speaker, not all plants need water. enriching (1) with the negation of (4) allows for the interpretation presented in (3), which we will refer to as an si, to emerge. to explain the underlying process through which sis arise, several theoretical accounts have been proposed. on the one hand, we find what can be called the enrichment approach: a family of accounts sharing the assumption that sis emerge via an extra-linguistic, context-dependent pragmatic process based on reasoning about the context, the speaker’s epistemic state, and the speaker’s intentions (sperber & wilson 1986, carston 2006, huang & snedeker 2009, geurts 2010). besides the enrichment approach, another prominent approach is the default approach (levinson 2000, chierchia 2004). substantial differences exist between accounts within this latter family. at the same time, these accounts share the assumption that sis do not necessarily involve a cognitive cost. rather, according to the default approach, sis are generated automatically on the basis of semantic rules (chierchia 2004) or heuristics of language (levinson 2000). the enrichment approach and the default approach offer different predictions regarding si generation and processing. according to the enrichment approach, sis are secondary interpretations (i.e., they are generated after the literal meaning of the sentence is computed) and require extra processing effort compared to literal meanings. according to the default approach, especially according to levinson’s (2000) proposal, sis are interpreted directly, and canceling the implicature in order to derive the literal interpretation of the sentence (“at least some and maybe all”) incurs a cognitive cost. a wealth of literature has been devoted to shed light on this debate. in the next section, we describe some of the relevant empirical evidence. 1.1. empirical evidence on si generation. in the typical si task (sentence evaluation task with underinformative sentences, see noveck 2001, bott & noveck 2004), participants are presented with some-underinformative sentences (e.g., some elephants have trunks). the rejection of such sentences implies that the si (corresponding to the “not all”-meaning of some) has been derived, because if one interprets the sentence as “not all elephants have trunks”, then the sentence is false. on the other hand, acceptance of some-underinformative sentences naturally follows from the derivation of the literal interpretation of the sentence. the available experimental evidence shows that typical adult language users can readily access both the literal interpretation of sentences and their sis. this is confirmed by the fact that adult individuals are easily capable of interpreting some with both its literal meaning (“at least some and maybe all”) and with its si-derived interpretation (“not all”) when instructed to do so in an experimental setting (bott & noveck 2004, exp. 1). at the same time, in experimental tasks in which adult individuals are not explicitly instructed to opt for one or the other interpretation, their interpretations distribute bimodally: in many experiments (e.g., bott & noveck 2004, exp. 3), roughly 40% of individuals consistently accept some-underinformative statements, thus opting for the literal interpretation; the remaining 60% of participants reject some-underinformative sentences, thus demonstrating a preference for the si-derived meaning. the observed variation in individuals’ tendency to generate sis has been attributed to various factors, including personality traits (feeney & bonnefon 2013) and cognitive skills (khorsheed & gotzner 2023). among those, working memory has emerged, at least in some studies, as a significant predictor of si generation tendencies (antoniou et al. 2016, nys et al. 2024, cf. heyman & schaeken 2015). proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 237 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ furthermore, when sis are derived in a sentence evaluation task, participants’ reaction times invariably emerge as significantly slower compared to the reaction times associated with the literal interpretation of different types of sentences (bott & noveck 2004, tomlinson et al. 2013, spychalska et al. 2016, van tiel et al. 2019, ronderos & noveck 2023). this slower processing is seen as an indication of an additional process taking place after the computation of the literal interpretation of the sentence. thus, a large body of literature supports the hypothesis that, as the enrichment approach suggests, sis are secondary interpretations and are associated with cognitively effortful processing. data from first language acquisition appear to point in the same direction. despite demonstrating good knowledge of the semantics of quantifiers, when presented with underinformative sentences like some elephants have trunks, 7to 11-year-old children tend to accept the statements (noveck 2001). likewise, when presented with a scenario in which all three out of three horses jump over a fence, 5-year-olds accept underinformative statements such as some of the horses jumped over the fence. results from both paradigms, therefore, indicate that children strongly prefer the literal meaning of some (“at least some and maybe all”). moreover, these results suggest that the ability to derive sis develops gradually in language acquisition, and is not yet fully adultlike even in adolescence (noveck 2001, porrini 2024). from the perspective of the enrichment approach, this may be unsurprising: if it is true that sis are effortful even for adults, albeit being within the reach of their cognitive skills, children’s still-developing cognitive abilities may explain why they do not access sis early in language acquisition and why they become adult-like only later, as their cognitive system matures. accounts attributing children’s difficulties with sis to limited cognitive skills or resources have been explicitly offered by reinhart (2004), pouscoulous et al. (2007), and mognon et al. (2021a, 2021b). second language (l2) learning represents another interesting testing ground for theories of sis and, in particular, for the hypothesized cognitive cost associated with si generation. being modulated by many factors, l2 processing appears generally less automatic, more effortful, and more reliant on cognitive resources such as working memory (reichle et al. 2016) than l1 processing. because of this, if deriving sis is cognitively effortful, it is expected that l2 speakers derive a reduced rate of sis compared to l1 speakers. in this regard, however, results are mixed: in some studies, no difference between l1 and l2 has been found (antoniou et al. 2019). in other studies, fewer sis were generated by l2 learners compared to l1 speakers (mazzaggio et al. 2021), and the si rate was found to be modulated by l2 proficiency (khorsheed et al. 2022). besides difficulties due to limited cognitive skills and resources, a reduced rate of sis could also stem from other causes. interestingly, in the realm of language acquisition, other explanations have been proposed for children’s non-adult-like si generation rates, a prominent one being the pragmatic tolerance account (katsos & smith 2010, katsos & bishop 2011, katsos 2014). as an account of children’s inferential skills, the pragmatic tolerance account is not often discussed in the literature on adult si generation. however, as we clarify below, we believe this account to be potentially relevant for explaining si generation in adult l2 speakers. in the remainder of the paper, we present an experimental study aimed at investigating the hypothesis that adult l2 speakers may derive fewer si interpretations compared to adult l1 speakers, not because of limited cognitive resources, but–as it has been argued for child (l1) speakers–for reasons related to pragmatic tolerance. 1.2. the pragmatic tolerance account. according to the pragmatic tolerance account (katsos & bishop 2011), children appear to generate fewer sis than adult (l1) speakers because, unlike adults, they are generally more tolerant towards pragmatic violations. the reasoning is as proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 238 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ follows: a sentence like some elephants have trunks is, in fact, not literally false, and children simply are more lenient than adults when judging this type of sentence. because of this, the methodology customarily used in si studies (in which participants are asked to either reject or accept some-underinformative sentences) misrepresents children’s actual pragmatic competence. however, when the appropriate methodology is used, children can perform adult-like, showing their ability to recognize the infelicity of some-underinformative sentences. in support of this claim, katsos & bishop (2011) designed a ternary truth-value judgment task. exactly as in binary truth-value judgment tasks, in ternary truth-value judgment tasks, the task requires evaluating different types of sentences. however, instead of binary response options (rejection vs. acceptance), participants are given a ternary scale including an intermediate option (corresponding to “true but sub-optimal”). katsos & bishop (2011) showed that, when asked to judge underinformative some-sentences (the mouse picked up some of the carrots in a context in which all five out of five carrots have been picked up by the mouse), children and adults performed alike, both overwhelmingly selecting the intermediate option. according to the authors, this result suggests that children are as sensitive to underinformativeness as adults. the difference between children and adults attested in the classical binary tasks is simply due to a different attitude towards pragmatic infelicity. importantly, according to katsos & bishop (2011), children’s disposition to tolerate underinformativeness may stem from their reduced language input. having had less exposure to language than adults, children may be less confident in their linguistic judgments. this would make them more tolerant and more prone to accept underinformative utterances. in light of this, could pragmatic tolerance also play a role in adult l2 speakers’ si generation? in other words, do l2 speakers accept underinformative sentences to a larger degree than adult l1 speakers do, and if so, can this be explained by their higher pragmatic tolerance due to their limited l2 exposure? in what follows, we present our study aimed at investigating these research questions. 2. current study 2.1. methods 2.1.1. participants. ninety-one adult speakers, all native speakers of dutch, participated in the experiment. participants were recruited via a local news bulletin in a small rural town in the netherlands. 2.1.2. design and procedure. the 91 participants were randomly divided in two groups: 43 participants were included in the l1 group and were tested in their l1 (dutch); 48 participants were included in the l2 group and were tested in their l2 (english). the si task was a ternary sentence evaluation task. in each trial, participants heard a recorded sentence and had to rate it as quickly as possible, choosing between three options: false, a bit true, true (for the l2 group) and onwaar, een beetje waar, waar (i.e., the dutch translations of the english terms, for the l1 group). following bott & noveck (2004), the task included six conditions (see table 1): the critical some-underinformative condition and five control conditions with the quantifiers some and all (some-true, some-false, all-true, all-false, all-falseabsurd). participants saw 6 items for the some-underinformative condition and 2 items for each of the control conditions. the total number of items seen by each participant was, therefore, 16. the experiment was run as an online questionnaire using the software qualtrics. to ensure that all participants would correctly understand the task, instructions were given to both groups in their proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 239 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ native language (dutch). during the experimental phase, the same sentences were presented in dutch to the l1 group, and in their english translation to the l2 group. after the si task, the l2 group was asked to report on their perceived english proficiency on a 7-point likert scale and on their usage of english in daily life (in reading, speaking, and listening). furthermore, the l2 group had their l2 english proficiency assessed by means of the standardized lexical test for advanced learners of english (lextale, lemhöfer & broersma 2012). condition example sentences expected response english (l2 group) dutch (l1 group) some-underinformative some cats are mammals. sommige katten zijn zoogdieren. a bit true (intermediate option) some-true some pets are dogs. sommige huisdieren zijn honden. true (acceptance) some-false some insects are lions. sommige insecten zijn leeuwen. false (rejection) all-true all lions are mammals. alle leeuwen zijn zoogdieren. true (acceptance) all-false all birds are chickens. alle vogels zijn kippen. false (rejection) all-falseabsurd all ducks are insects. alle eenden zijn insecten. false (rejection) table 1: overview of the materials. the expected responses are based on the pragmatic tolerance account 3. results 3.1. l2 proficiency. lextale scores, calculated using the method recommended by lemhöfer & broersma (2012), range from 0 to 100. table 2 presents the summary statistics for our participants (l2 group only), while figure 1 visually displays the individual scores. these scores correspond to different levels of the common european framework (cef) for language levels (council of europe 2001). as reported in lemhöfer & broersma (2012), a lextale score below 60 corresponds to the cef level b1 or lower (lower intermediate user and lower), whereas a score above 80 corresponds to c1/c2 (advanced/proficient user). in light of the wide range of proficiency levels among our participants, we divided the l2 group into two subgroups using a median split. participants with lextale scores of 70 or below were classified as l2 low proficiency (n = 25 participants), and those with scores above 70 were classified as l2 high proficiency (n = 23 participants). range mean median 53.75 92.50 71.98 70.00 table 2: summary statistics of lextale scores (l2 group, n = 48) proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 240 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: lexttale score distribution (l2 group, each dot presenting a participant) participants’ self-reported usage (in the domains of reading, listening, and speaking) and self-rated proficiency showed only weak correlations with the lextale score (τb < .34). compared to selfratings, the lextale task has been shown to provide a finer-grained measure of english proficiency (lemhöfer & broersma 2012). thus, we decided to focus exclusively on lextale scores and not to consider the other measures further. 3.2. ternary sentence evaluation task. nine participants (8 from the l2 low proficiency group and 1 from the l2 high proficiency group) gave fewer than 80% correct responses in the control conditions. therefore, these participants were excluded from further analysis. after excluding these participants, performance on the control conditions was as expected: false sentences were overwhelmingly rejected, true sentences were overwhelmingly accepted, and the intermediate option was hardly ever selected (fig. 2). the control conditions were, therefore, not analyzed further. participants’ responses in the critical condition some-underinformative are shown in figure 3. figure 2: responses (%) on the five control conditions by all participants (n = 82) proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 241 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 3: responses (%) on the some-underinformative condition by the three participant groups (l1 vs. l2 high vs. l2 low) (n = 82) data analysis for the critical some-underinformative condition was conducted in r (r core team 2024, r version 4.2.3) and aimed to address three primary objectives. our first objective was to investigate whether the three groups responded differently when presented with a ternary scale. therefore, given the ordered nature of our response variable (true, a bit true, false), we modeled the cumulative probability of the response falling into a particular level or lower, based on the proficiency group (l1, l2 high proficiency, l2 low proficiency). our second objective was to examine whether participants’ tendency to interpret sentences literally (as indicated by their full acceptance of some-underinformative sentences) varied across proficiency groups. to address this issue, we analyzed our data focusing on the response true. our third objective was to assess whether the probability of displaying pragmatic tolerance (selecting the intermediate option in response to underinformativeness) was influenced by the proficiency group. to address this issue, we focused on the response a bit true. 3.2.1. responses on the ternary scale. first, to determine whether the three participant groups responded differently in the ternary judgment task, we analyzed our data using ordinal mixedeffect regression (r package ordinal, christensen 2019). as outcome variable, we included in our model the ordered variable response (true, a bit true, false) and as predictor the categorical variable proficiency with three levels: l1, l2 high proficiency, and l2 low proficiency. because the model violated the proportional odds assumption (assessed via harrell’s 2001 graphical method), we re-ran the analysis using the function clmm2 instead of clmm. this allowed for the inclusion of a scale effect for the predictor variable proficiency. we also included a by-participant random effect (note that clmm2 does not allow for more than one random effect and thus the by-item random effect was not included). the predictor proficiency turned out not to be significant (𝛽highproficiency = 0.25, p > .5; 𝛽lowproficiency = 1.19, p > .5), indicating that the three groups did not respond significantly differently when selecting one of the three response options (acceptance, intermediate option, rejection) in their judgment of some-underinformative sentences. 3.2.2. literal interpretations of underinformative sentences. our second objective was to determine whether the three participant groups showed a different tendency to interpret sentences literally by opting for the acceptance option (true) as opposed to the intermediate (a bit true) or the rejection (false) options. therefore, we dichotomized the responses by creating a binary outcome variable with two levels: acceptance (true responses) vs. other responses (a bit true + false responses together). we then used binary logistic regression modeling (r package lme4, proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 242 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ bates et al. 2015) to predict the likelihood of selecting true vs. other responses based on the proficiency group (l1 vs. l2 high proficiency vs. l2 low proficiency). our model, therefore, included proficiency as the main predictor and the maximum random effect structure allowed by our data (a by-participant random slope for proficiency and a by-trial random effect). no difference was found between our reference level (the l1 group) and the l2 high proficiency group (𝛽highproficiency = -0.09, p = .90). however, a difference emerged between the reference level (the l1 group) and the l2 low proficiency group, such that the latter group was significantly more likely to select the true response (𝛽lowproficiency = 1.25, p =.03). 3.2.3. intermediate responses to underinformative sentences. for our third objective, we ran a similar logistic regression analysis to assess participants’ tendency to select the intermediate response option. thus, we further dichotomized the responses by creating another outcome variable with two levels: intermediate option (a bit true responses) vs. other responses (true + false responses together). this logistic regression analysis was used to predict the likelihood of selecting a bit true vs. one of the other responses based on the proficiency group (l1 vs. l2 high proficiency vs. l2 low proficiency). thus, we included proficiency as main predictor and a by-participant random effect (because of convergence issues, no slopes and no by-trial random effects were included). the model showed that the intermediate option was less likely to be selected by the l1 group (our reference level) compared to the other two responses (𝛽 = -3.7, p < .001). importantly, the predictor proficiency was not significant (𝛽highproficiency = 0.73, p = .43; 𝛽lowproficiency = 0.81, p = .42). this indicates that the three groups did not differ in their preferences towards the intermediate option as opposed to the other choices. in summary, our ordinal regression analysis suggested that the three groups did not use the ternary scale significantly differently when judging some-underinformative sentences. the binary logistic regression analyses provided us with two additional findings: first, the probability of selecting the acceptance option was modulated by language proficiency; second, the probability of selecting the intermediate option was equally low in the three groups. 4. discussion and conclusions. with this study, we aimed to investigate the hypothesis that l2 speakers may generate fewer sis compared to l1 speakers because of a higher tolerance towards pragmatic violations, possibly stemming from their reduced language exposure compared to l1 speakers. following the pragmatic tolerance account, we expected l1 and l2 speakers to behave alike in our ternary sentence evaluation task and to opt for the intermediate option (a bit true) when judging some-underinformative sentences. contrary to our expectations, our l2 participants did not prefer to judge these sentences as “true but sub-optimal” using the intermediate option of the ternary scale. this is in contrast with the hypothesis that l2 speakers might be sensitive to underinformativeness and accept some-underinformative sentences in binary tasks only because they are more tolerant towards pragmatic infelicity. in other words, our results do not support the idea that the pragmatic tolerance account can be extended to l2 learning. importantly, however, if the pragmatic tolerance account provides a viable account of si generation in experimental tasks and if adult l1 speakers are fully capable of recognizing the infelicity of some-underinformative sentences, we would also expect l1 control speakers to preferentially select the intermediate option when judging underinformative sentences like some cats are mammals. this, however, was not the case. like the l2 speakers, the l1 speakers did not prefer the intermediate option; in fact, they selected the intermediate option in only 15% of the cases. proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 243 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ our findings, we believe, underscore three fundamental issues related to, (1) the role of language proficiency in si generation, (2) the cognitive costs involved in si generation and the conflicting findings emerging in previous l2 literature, and (3) the reliability of ternary judgment tasks for the study of pragmatics. 4.1. scalar implicature generation and language proficiency. the first issue warranting discussion concerns the fact that, in spite of the unexpected results in relation to the intermediate response in our study, a significant difference between groups did emerge in our data and was linked to language proficiency. the l2 low proficiency group tended to fully accept some-underinformative sentences significantly more often than the l1 group; in contrast, the l2 high proficiency group did not differ from the l1 group. on the one hand, the fact that the l2 high proficiency group showed a similar pattern of responses compared to the l1 group is not surprising, as the mastery of english of the l2 high proficiency group, based on the scores of the lextale task, was extremely high (equal to or above level c1 for most of the participants). the dutch participants in the l2 high proficiency group, in essence, were almost native-like in their l2 english. on the other hand, the l2 low proficiency participants accepted some-underinformative sentences more than the other two groups, thus displaying a stronger preference for the literal interpretation of such sentences. these results, we believe, align with recent findings showing the modulating effect of language proficiency on the tendency to accept the literal meaning of utterances, as opposed to generating sis (khorsheed et al. 2022, mazzaggio et al. 2021). given that our results do not support the pragmatic tolerance account, how can we explain the modulating role of language proficiency on the tendency to interpret sentences literally or pragmatically? as mentioned in the introduction, according to the default approach to sis, accepting the literal interpretation of some-underinformative sentences should come at a cost. our results do not support this idea, given that the only group that preferentially accepted some-underinformative sentences in our experiment was the low proficiency group, that is to say, the group of participants whose cognitive resources were arguably more limited. on the other hand, according to the enrichment approach, sis should incur a cognitive cost. our results support this hypothesis: participants with fewer cognitive resources available (i.e., participants in the l2 low proficiency group) opted for the cognitively less demanding option, that is to say, they interpreted some-underinformative sentences literally. 4.2. the cognitive cost of scalar implicature generation. as a second issue, we would like to discuss the cognitive costs of si generation in the l2 and the mixed findings of previous literature. previous studies on sis in the l2 do not unequivocally show that l2 speakers generate fewer sis compared to l1 speakers. in some experiments, no difference emerged between l1 and l2 groups (e.g., antoniou & katsos 2017). on the other hand, slabakova (2010) found that l2 speakers are more likely to generate sis than l1 speakers. how to explain the mixed evidence? one important factor seems to be language immersion. in the study of slabakova (2010), at the time of testing all l2 participants were immersed in an l2 environment, and were attending university courses in their l2. arguably, therefore, slabakova’s (2010) participants were accustomed in their daily lives to invest extra cognitive resources in l2 processing. this factor may have made them prone to derive sis, even more than the l1 speakers included in slabakova’s (2010) study. in line with this, as already noted by mazzaggio et al. (2021), immersion in the l1 is likely to give rise to the opposite effect: both in mazzaggio et al. (2021) and in khorsheed et al. (2022), participants were immersed in their l1 environment, and a reduced rate of sis was found in the l2 compared proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 244 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ to the l1. our study confirms this pattern of results by finding a reduced rate of sis in a group of l2 low proficiency participants who were also immersed in their l1 environment. besides language immersion, a second factor that could explain the mixed findings in the literature is depth of processing. as suggested by mazzaggio et al. (2021) and discussed by khorsheed & van tiel (2024), the difficulties generally experienced by l2 speakers in si generation may disappear when participants are given the chance to engage in deeper processing. in experimental paradigms in which participants are put under time constraints or stimuli are presented auditorily, l2 speakers struggle to generate sis because si generation imposes a burden on their already reduced cognitive resources. in line with this hypothesis, khorsheed & van tiel (2024) demonstrated experimentally that when the task grants participants more time to process the sentences, l2 speakers are able to overcome the difficulties connected to si generation and show native-like performance. again, our results are in line with these observations. in our experiment, in which the stimuli were presented auditorily and participants were asked to answer as quickly as possible, l2 speakers (at least those with lower proficiency levels) were more likely to fully accept some-underinformative sentences. arguably, this is because they did not have sufficient time to engage in deeper processing and to go beyond the literal meaning of the utterances. 4.3. ternary judgment tasks. finally, our study suggests that the reliability of ternary judgment tasks for gauging the pragmatic-inferential skills of different populations should not be taken for granted. in fact, our adult l1 speakers, despite presumably being able to recognize underinformativeness, failed to select the intermediate response when judging some-underinformative sentences. how can we explain their unexpected behavior? could these puzzling results be due to a flaw in our experimental design? our study is not the first si experiment based on a ternary judgment task in which an unexpected pattern of results emerged. wampers et al. (2018), for instance, tested the si generation skills of patients with psychosis and adult l1 controls. the experiment included a task that was virtually identical to the one used in our experiment, namely a sentence evaluation task with underinformative sentences and patently false and true control sentences with some and all. the adult l1 control group in wampers et al. (2018) preferentially rejected some-underinformative sentences, and both groups failed to show a preference for the intermediate option. intermediate answers were only given in 22% and 28% of the cases by patients and controls, respectively. another study in which the intermediate option was hardly ever selected by participants is schaeken et al. (2021). in this study, using a different design and a 5-point likert scale, the authors found that the selection rate of the intermediate option as a response to some-underinformative sentences was invariably below or equal to 15%. this was true both for the control group (adult l1 language users) and for the group of patients under investigation (i.e., individuals with schizophrenia and other psychotic disorders). this pattern of results (i.e., no preference for the intermediate option) is surprising if one considers that in the original experiment of katsos & bishop (2011), the control group of adults was reported to “invariably” select intermediate responses, suggesting a rate of selection of intermediate responses close to 100%. at this point, it is worth highlighting an important difference in experimental design between the study of katsos & bishop (2011) and the other studies mentioned above. whereas, in line with wampers et al.’s (2018) study, we used the classical sentence evaluation task (requiring participants to judge sentences like some cats are mammals in isolation, without visual context), katsos and bishop (2011) adopted a truth-value judgment task (requiring participants to consider, for instance, a visual scene in which a mouse picked up five out of five carrots and judge the sentence the mouse picked up some of the carrots). speculatively, we suggest that the inclusion of a visual proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 245 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ context in the task motivates participants to evaluate the visual context instead of the sentence. in truth-value judgment tasks, selecting the intermediate option could be the participants’ strategy to suggest that a small modification of the visual context (e.g., the mouse picking up four instead of five carrots) could suffice to make the sentence felicitous in that context. the same reasoning is, however, not applicable to sentence evaluation tasks: the very meaning of the words cats and mammals, for instance, is semantically determined and hence no small modification to the test materials suffices to make the sentence felicitous. be as it may, the fact that a preference for the intermediate option cannot be replicated in experiments using different tasks casts doubts on the primary assumption of the pragmatic tolerance account. that is, if language users are sensitive to underinformativeness (which is a prerequisite for si generation), they are assumed to select the intermediate option to signal their sensitivity towards underinformativeness, while rejecting patently false utterances and accepting patently true utterances. therefore, we argue that caution should be exercised when using ternary judgment tasks. in particular, the idea that ternary judgment tasks can be used, as is sometimes claimed, as a finer measure of the pragmatic skills of various populations deserves to be reconsidered. in summary, our study does not bring support to the idea that the pragmatic tolerance account can be extended to si generation in the l2. however, our findings add to previous research by suggesting that sis are effortful. in adult l2 speakers, this effort leads to a greater preference for the literal interpretation of sentences when the task does not grant participants sufficient time to engage in deeper processing and when the l2 proficiency is relatively low. references antoniou, kyriakos, chris cummins & napoleon katsos. 2016. why only some adults reject under-informative utterances. journal of pragmatics 99. 78–95. https://doi.org/10.1016/j.pragma.2016.05.001. antoniou, kyriakos & napoleon katsos. 2017. the effect of childhood multilingualism and bilectalism on implicature understanding. applied psycholinguistics 38(4). 787–833. https://doi.org/10.1017/s014271641600045x. antoniou, kyriakos, alma veenstra, mikhail kissine & napoleon katsos. 2019. how does childhood bilingualism and bi-dialectalism affect the interpretation and processing of pragmatic meanings? bilingualism: language and cognition 23(1). 186–203. https://doi.org/10.1017/s1366728918001189. bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67(1). 1–48. https://doi.org/10.18637/jss.v067.i01. bott, lewis & ira a. noveck. 2004. some utterances are underinformative: the onset and time course of scalar inferences. journal of memory and language 51(3). 437–457. https://doi.org/10.1016/j.jml.2004.05.006. carston, robyn. 2006. relevance theory and the saying/implicating distinction. in: laurence r. horn & gregory ward (eds.), the handbook of pragmatics. 633–656. 1st edn. wiley. https://doi.org/10.1002/9780470756959.ch28. chierchia, gennaro. 2004. scalar implicatures, polarity phenomena and the syntax/pragmatics interface. in: adriana belletti (ed.), structures and beyond. oxford: oxford university press. 39–103. proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 246 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ christensen, rune haubo b. 2019. ordinal—regression models for ordinal data. r package version 2019.12-10. https://cran.r-project.org/package=ordinal. council of europe. 2001. common european framework of reference for languages: learning, teaching, assessment. cambridge: cambridge university press. feeney, aidan & jean-françois bonnefon. 2013. politeness and honesty contribute additively to the interpretation of scalar expressions. journal of language and social psychology 32(2). 181–190. https://doi.org/10.1177/0261927x12456840. geurts, bart. 2010. quantity implicatures. cambridge: cambridge university press. https://doi.org/10.1017/cbo9780511975158. grice, herbert paul. 1989. studies in the way of words. cambridge: harvard university press. harrell, frank e. 2001. regression modeling strategies: with applications to linear models, logistic regression, and survival analysis. new york: springer. heyman, tom & walter schaeken. 2015. some diferences in some: examining variability in the interpretation of scalars using latent class analysis. psychologica belgica 55(1). 1–18. https://doi.org/10.5334/pb.bc. huang, yi ting & jesse snedeker. 2009. online interpretation of scalar quantifiers: insight into the semantics–pragmatics interface. cognitive psychology 58(3). 376–415. https://doi.org/10.1016/j.cogpsych.2008.09.001. katsos, napoleon. 2014. scalar implicature. in: danielle matthews (ed.), pragmatic development in first language acquisition. 183–198. amsterdam: john benjamins. katsos, napoleon & dorothy v.m. bishop. 2011. pragmatic tolerance: implications for the acquisition of informativeness and implicature. cognition 120(1). 67–81. https://doi.org/10.1016/j.cognition.2011.02.015. katsos, napoleon & nafsika smith. 2010. pragmatic tolerance or a speaker-comprehender asymmetry in the acquisition of informativeness? in: katie franich, kate m. iserman, and lauren l. keil (eds.), proceedings of the 34th annual boston university conference on language development, vols 1 and 2. 221–232. khorsheed, ahmed & nicole gotzner. 2023. a closer look at the sources of variability in scalar implicature derivation: a review. frontiers in communication 8. https://doi.org/10.3389/fcomm.2023.1187970. khorsheed, ahmed & bob van tiel. 2024. why second-language speakers sometimes, but not always, derive scalar inferences like first-language speakers: effects of task demands. language acquisition 1–19. https://doi.org/10.1080/10489223.2024.2383574. khorsheed, ahmed, sabariah md. rashid, vahid nimehchisalem, lee geok imm, jessica price & camilo r. ronderos. 2022. what second-language speakers can tell us about pragmatic processing. plos one 17(2). e0263724. https://doi.org/10.1371/journal.pone.0263724. lemhöfer, kristin & mirjam broersma. 2012. introducing lextale: a quick and valid lexical test for advanced learners of english. behavior research methods 44. 325–343. levinson, stephen c. 2000. presumptive meanings: the theory of generalized conversational implicature. cambridge, ma: mit press. mazzaggio, greta, daniele panizza & luca surian. 2021. on the interpretation of scalar implicatures in first and second language. journal of pragmatics 171. 62–75. https://doi.org/10.1016/j.pragma.2020.10.005. mognon, irene, simone a. sprenger, sanne j. m. kuijper & petra hendriks. 2021a. complex inferential processes are needed for implicature comprehension, but not for implicature production. frontiers in psychology 11. 1–18. https://doi.org/10.3389/fpsyg.2020.556667. proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 247 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ mognon, irene, simone sprenger, sanne kuijper & petra hendriks. 2021b. balancing the (horn) scale: explaining the production-comprehension asymmetry for scalar implicatures. in: alexandros kalomoiros & lefteris paparounas (eds.), upenn working papers in linguistics 27(1). philadelphia, pa: university of pennsylvania. noveck, ira a. 2001. when children are more logical than adults: experimental investigations of scalar implicature. cognition 78(2). 165–188. https://doi.org/10.1016/s0010-0277(00)001141. noveck, ira a. 2018. experimental pragmatics: the making of a cognitive science. cambridge: cambridge university press. nys, bojan luc, wai wong & walter schaeken. 2024. some scales require cognitive effort: a systematic review on the role of working memory in scalar implicature derivation. cognition 242. 105623. https://doi.org/10.1016/j.cognition.2023.105623. porrini, anna teresa. 2024. cooperative intentions and epistemic reasoning in scalar implicature derivation: a developmental perspective. trento, italy: università degli studi di trento dissertation. https://hdl.handle.net/11572/410450. pouscoulous, nausicaa, ira a. noveck, guy politzer & anne bastide. 2007. a developmental investigation of processing costs in implicature production. language acquisition 14(4). 347– 375. https://doi.org/10.1080/10489220701600457. r core team. 2024. r: a language and environment for statistical computing. vienna, austria: r foundation for statistical computing. https://www.r-project.org/. reichle, robert v., annie tremblay & caitlin coughlin. 2016. working memory capacity in l2 processing. probus 28(1). 29–55. https://doi.org/10.1515/probus-2016-0003. reinhart, tanya. 2004. the processing cost of reference set computation: acquisition of stress shift and focus. language acquisition 12(2). 109–155. https://doi.org/10.1207/s15327817la1202_1. ronderos, camilo r. & ira noveck. 2023. slowdowns in scalar implicature processing: isolating the intention-reading costs in the bott & noveck task. cognition 238. https://doi.org/10.1016/j.cognition.2023.105480. schaeken, walter, linde van de weyer, marc de hert & martien wampers. 2021. the role of working memory in the processing of scalar implicatures of patients with schizophrenia spectrum and other psychotic disorders. frontiers in psychology 12. 1447. https://doi.org/10.3389/fpsyg.2021.635724. slabakova, roumyana. 2010. scalar implicatures in second language acquisition. lingua 120(10). 2444–2462. https://doi.org/10.1016/j.lingua.2009.06.005. sperber, dan & deirdre wilson. 1986. relevance: communication and cognition. vol. 142. cambridge, ma: harvard university press. spychalska, maria, jarmo kontinen & markus werning. 2016. investigating scalar implicatures in a truth-value judgement task: evidence from event-related brain potentials. language, cognition and neuroscience 31(6). 817–840. https://doi.org/10.1080/23273798.2016.1161806. tiel, bob van, elizabeth pankratz & chao sun. 2019. scales and scalarity: processing scalar inferences. journal of memory and language 105. 93–107. https://doi.org/10.1016/j.jml.2018.12.002. tomlinson, john m., todd m. bailey & lewis bott. 2013. possibly all of that and then some: scalar implicatures are understood in two steps. journal of memory and language 69(1). 18– 35. https://doi.org/10.1016/j.jml.2013.02.003. proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 248 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ wampers, martien, sofie schrauwen, marc de hert, leen gielen & walter schaeken. 2018. patients with psychosis struggle with scalar implicatures. schizophrenia research 195. 97–102. https://doi.org/10.1016/j.schres.2017.08.053. proceedings of elm 3: 236-249, 2025 irene mognon, amber l. marree, and petra hendriks: are second language speakers more pragmatically tolerant?. 249 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ experimental paradigms on scalar implicature estimation zhuang qiu, casey d. felton, zachary n. houghton & masoud jasbi* abstract. experimental research on the processing of scalar implicatures (sis) relies on behavioral tasks that purport to measure the rate at which scalar implicatures are computed within an experimental paradigm. two paradigms, the truth value judgment task (tvjt) (gordon, 1998, crain & thornton, 2000) and the picture selection task (pst) (gerken & shady, 1998) have dominated the experimental pragmatics literature; yet the effects of task choice on implicature rate have remained underexplored. here we report the results of three studies testing participants in a tvjt and a pst using three different linguistic scales in english: “or-and”, “someall”, and “ad-hoc”. we varied the task (tvjt vs. pst) within subjects in the first experiment and between subjects in the second. the third experiment examined a variant of the pst called the hidden card task (hct) which is increasingly used in the context of priming research (bott & chemla, 2016). we found that the estimated rate of scalar implicature computation varied noticeably between different tasks as well as scales. this suggests that the experimental paradigm itself has a significant impact on our estimates of the implicature rate for a given linguistic scale, and thus, researchers studying scalar implicatures need to carefully consider the pragmatics of the task itself when designing experimental studies and interpreting their results. keywords. scalar implicature; implicature computation; truth value judgment task; picture selection task; experimental pragmatics 1. introduction. an intriguing feature of human language is the ability to enrich the literal meanings of utterances with pragmatic implicatures (grice, 1975, horn, 1972, gazdar, 1979, hirschberg, 1985, levinson, 1983, 2000), as shown in the following examples: (1) a. “some students were late today.” implicature: not all students were late today. b. “sam had a hot dog or a hamburger for lunch.” implicature: sam did not have a hot dog and a hamburger for lunch. c. “alice bought a shirt from the store.” implicature: from that store, alice only bought a shirt but nothing else. according to standard accounts (horn, 1972, gazdar, 1979, levinson, 1983), word pairs such as “some-all” and “or-and” form a scale of increasing informativeness with a unidirectional entailment relation. if all students were late today, it must be the case that some students were late, but not vice versa. therefore, the quantifier “all” is more informative than “some”. similarly, “and” is argued to be more informative than “or” because if sam had a hot dog and a hamburger for lunch, it must be the case that he had a hot dog or a hamburger for lunch, but not vice versa. the less informative item on the scale is semantically compatible with the cases in which the more informative item holds true; however, the assertion of the lower item implies the negation of the higher item. the computation of such implicature is believed to be governed by general principles of conversation and involves reasoning about the possible alternatives that the * zhuang qiu (zhuangqiu@cityu.edu.mo), city university of macau; casey d. felton (cdfelton@ucdavis.edu), zachary n. houghton (znhoughton@ucdavis.edu), and masoud jasbi (jasbi@ucdavis.edu), university of california, davis. proceedings of elm 3: 308-318, 2025 c©2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi published by the lsa with permission of the author(s) under a cc by license. 308 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ speaker could have said (grice, 1975, 1978). for example, interlocutors are expected to make their words as informative as required while being truthful at the same time. abiding by these principles, when hearing the utterance “some students were late today”, the listener reasons that the speaker did not use a more informative alternative “all students were late today” because that alternative is not true. this gives rise to the scalar implicature (si) that the use of “some” implies the negation of “all” (see gamut, 1991: 207 for a similar analysis of the “or-and” scale). the same reasoning process also underlies (1c), which is argued to involve the scale of a setsubset relationship (hirschberg, 1985:109). in this case, the item “shirt” is a subset of all the items that can be purchased from that store in that conversational context. this type of scale is called an “ad hoc” scale because the implicature is derived on an ad hoc basis depending on the context (bott & chemla, 2016). more recent theoretical frameworks either attribute sis to syntactic operations at the word level (chierchia, fox, & spector 2012, chierchia, 2013) or highlight conversational contexts rather than lexical scales (sperber & wilson, 1995; degen & tanenhaus, 2015, 2019). these accounts differ with regard to the mechanisms that derive sis. however, all of these accounts acknowledge the distinction between upper-bounded inferences (e.g. “some” as “some but not all”) and more literal, lower-bounded interpretations (e.g. “some” as “at least one”) in expressions containing scalar items as shown in (1a-b). there is an increasing number of empirical studies on the processing and acquisition of sis in the relatively new field of experimental pragmatics. such studies rely on behavioral tasks that aim to measure the rate at which the upper-bounded inferences are computed within an experimental paradigm, among which the truth value judgment task (tvjt) (gordon, 1998, crain & thornton, 2000) and the picture selection task (pst) (gerken & shady, 1998) have dominated the experimental pragmatics literature (as for other widely cited paradigms, see huang & snedeker, 2009, 2011, grodner et al., 2010 for the visual world paradigm, also see degen & tanenhaus, 2015 for the gumball paradigm). in a typical tvjt study, participants are required to judge whether a sentence is true or false based on some background information (the world knowledge shared across the participants or other information provided prior to the target sentence). in the critical trials, the sentence to be judged is pragmatically infelicitous given the background information, but remains logically true (e.g., “some elephants have trunks”). therefore, if a participant responds with “false”, the experimenter concludes that they have computed an implicature, but not if they respond with “true”. in a pst, participants are required to select a picture that best matches a given sentence from a set of pictures. in the critical trials, the sentence is logically compatible with more than one picture, but the implicature of the sentence only matches one picture. for example, the sentence “there is a cat in the picture” is logically compatible with 1) a picture with only a cat and 2) a picture with a cat as well as a dog, but the pragmatic interpretation of the sentence is only compatible with the first picture. therefore, if a participant selects the first picture, the experimenter concludes that they have computed an implicature. recent priming research on scalar implicature (rees, carter & bott, 2023, rees & bott, 2018, bott & chemla, 2016) adopted a modified version of the pst in which the pragmatically felicitous card (e.g. the card with only a cat in the previous case) is replaced by a card with the text “better picture?” on it, and participants were instructed to select the “better picture” card if they feel the content of the other card is not optimally described by the given sentence. in this case, the card whose content is fully visible to the participants is the one which is logically compatible with the given sentence but pragmatically infelicitous. participants’ choice between the “better picture” card and the fully visible card is interpreted as the presence or absence of a scaproceedings of elm 3: 308-318, 2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi: experimental paradigms on scalar implicature estimation. 309 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ lar implicature. this version of the pst is referred to as the hidden card task (hct) since the hypothetical content of the “better picture” card is hidden from participants. both the tvjt and the pst have been adopted as a measure of si computation to study topics including but not limited to (1) the debate over the psychological nature of sis as either default and automatic or as secondary and effortful (noveck & posada, 2003, bott & noveck, 2004, de neys & schaeken, 2007), (2) factors that affect whether a si is computed in a communicative context (geurts & pouscoulous 2008, 2009; chemla & spector, 2011, potts et al., 2015), (3) variations in si computation across different lexical scales (doran et al., 2012, van tiel et al., 2016), and (4) the development of pragmatic knowledge in the first and second language context (papafragou & musolino, 2003, katsos & bishop, 2011, horowitz, schneider & frank, 2018, slabakova, 2010, feng & cho, 2019). there seems to be a tacit assumption that different versions of the tvjt and the pst paradigm are variations of objective measures of si computation, a construct whose psychological nature and structural properties are still debated. this assumption is problematic as there has been increasing concern that the experimental paradigm itself has a significant impact on the estimates of the implicature rate. papafragou and musolino (2003) showed that children in general did not judge the underinformative descriptions as “false” in a tvjt task, though they knew that those descriptions were not optimal. katsos and bishop (2011) argued that both adults and children are more tolerant to pragmatic infelicity than logical violations, and this explains why underinformative statements tend not to be judged as “false”. crucially, the tolerance to pragmatic infelicity is not equal to the absence of pragmatic inferences, but these two aspects tend to be confused in a binary tvjt task. thus, katsos and bishop included a pst in their study as a complement to tvjt. since in a pst, participants were asked to select the most felicitous interpretation rather than judge the absolute truth of an interpretation, the “pragmatic tolerance cannot cloud the interpretation of the participants’ performance” (2011:14). katsos and bishop brought to the foreground questions such as whether the observed si rate is contingent on the task choice between tvjt and pst, and if so, how these two paradigms differ in measuring si computation. however, there has been limited research that compares tvjt and pst in controlled settings using the same set of experimental items with the same group of participants. we conducted three experiments to explore the effect of task variation on the estimated rate of scalar implicature computation. specifically, we compared the tvjt, pst, and hct regarding the estimated rate of scalar implicature computation, focusing on the following research questions: 1. when a tvjt, pst, and hct consisting of theoretically parallel items are administered to the same group of participants, will the observed si rate differ depending on the task type? 2. how much variation is there in the observed si computation across different scale types? is such variation modulated by the task type? 3. how reliable are the tvjt, pst, and hct as measures of si computation? the first two experiments followed a two by three design in which the task type (tvjt vs. pst) and scale type (“some-all”, “or-and” and “ad-hoc”) were independent variables and the answers elicited were the dependent variable. in the first experiment, both task type and the scale type were manipulated within participants, while the second experiment treated task type as a between participants variable. the third experiment adopted the hct with experimental items paralleling those of the first two experiments. we found that the reliability of all three tasks are high, but the estimated rate of scalar implicature computation varied noticeably between tasks and scales. proceedings of elm 3: 308-318, 2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi: experimental paradigms on scalar implicature estimation. 310 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 2. experiment 1. the goal of this experiment was to test the same participants in both the tvjt and the pst tasks. we used three different scales: “or-and”, “some-all”, and ad hoc. the withinparticipant design of this study allowed us to see how much the same individual can vary in their responses to each task depending on the linguistic scale. 2.1. methods. fifty participants were recruited from the online crowdsourcing platform prolific. they were instructed to answer a series of questions corresponding to both a tvjt and a pst in a single qualtrics survey. in each tvjt trial, participants were presented with a card that had one or more animal images on it and a sentence describing the content of the card. they were instructed to rate the sentence as either true or false. in the critical trials (figure 1a), the description was logically true but pragmatically infelicitous, and participants’ judgment was coded as whether or not a si was computed. in each pst trial, participants were presented with two cards and a sentence describing the contents of at least one card. participants were instructed to choose the card that best matched the sentence. in the critical pst trials (figure 1b), the sentence was logically compatible with both cards, but the implicature of the sentence only matched one card, and thus, participants’ preference of one card over the other was coded as whether or not a si was computed. figure 1: an example of a critical item in tvjt (1a), pst (1b), and hct (1c). the response showing the computation of si for each task is marked by the rectangle. this example concerns the “or-and” scale, while other experimental items may use the “some-all” or the “ad-hoc” scale. in addition to the images of cats and dogs, images of elephants were also used in the design of the cards. the position of the two cards in (1b) and (1c) was randomized in the experiment. task types “tvjt” and “pst” pertain to experiment 1 and 2, while “hct” is for experiment 3. critical trials in the tvjt paralleled those in the pst as they used the same picture and the same stimuli sentence. stimuli sentences came from three different scales to generate sis: the “some-all” scale, the “or-and” scale, and the “ad-hoc” scale (bott and chemla, 2016). in total, there were nine critical trials for each task, and each critical trial appeared twice in the experiment. in addition to the critical trials, 81 control trials were created in the form of either tvjt or pst to check participants’ engagement in the task or to provide information about participants’ baseline preference (table 1). in the experiment, critical trials and a subset of control trials (40 proceedings of elm 3: 308-318, 2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi: experimental paradigms on scalar implicature estimation. 311 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ out of 81) were blocked, and trials in each block were randomly presented to participants. the control block was always presented after the experimental block, in order to make sure scalar inferences are not affected by the lexical choices in the control trials. 2.2. analysis. of the 50 participants recruited, seven were excluded for low accuracy on the attention checking control items, leaving 43 participants in the final analysis. for the purpose of this study, we only analyzed the critical (experimental) trials in the tvjt and the pst. table 1: experimental manipulations using stimuli from the “some-all” scale as an example. task types “tvjt” and “pst” pertain to experiment 1 and 2, while “hct” is for experiment 3. we constructed a bayesian logistic generalized linear model (bürkner, 2017) to explore how task variation influences si computation. the probability of computing sis was modeled as a function of task type (tvjt vs pst), scale (“some-all”, “or-and”, and “ad hoc”) and their interactions. the model also included by-subject random intercepts, scale-by-subject slopes, task-bysubject slopes, and slopes for the interaction of scale and task by subject. the predictors were dummy coded. since each critical item appeared twice in the experiment, we also constructed bayesian logistic generalized linear models to check if the si rate differed depending on whether the participants encountered the item for the first or the second time. we separated tvjt trials from pst trials, and for each of the tasks, the probability of computing sis was modeled as a function of trial iteration (first iteration or the second iteration). maximal random effects structures were constructed including subject and item intercepts and slopes (barr et al., 2013). we then compared the composite reliability of tvjt and pst based on the coefficient alpha calculated following the kuder-richardson formula 20 (cronbach, 1951, kuder & richardson, 1937). 2.3. results. we found main effects of task type, scale, and their interactions on the estimated rate of si computation (figure 2). compared with the baseline “or-and” trials (in pst) participants computed more sis in “some-all” trials (beta = 17.82, ci = [11.34, 27.79]) and “ad hoc” trials (beta = 14.12, ci = [9.11, 21.53]). for the “or-and” trials, the rate of computing sis in pst (baseline) was the same as that in tvjt (beta = -1.73, ci = [-4.82, 0.98]); however, for the “some-all” trials and “ad hoc” trials, the rates of computing sis were significantly decreased in the tvjt (beta = -19, ci = [-31.66, -11.26]; beta = -20.84, ci = [-34.98, -12.46], respectively). moreover, there was no effect of item iteration on the probability of computing sis for both proceedings of elm 3: 308-318, 2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi: experimental paradigms on scalar implicature estimation. 312 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ tvjt (beta = 0.15, ci = [-0.91, 1.07]) and pst (beta = 0.42, ci = [-0.38, 1.34]). participants provided the same answer to the same question regardless of whether they saw it the first time or the second time. the coefficient alpha for the 18 tvjt and 18 pst critical items was 0.92 and 0.83, respectively. the observed differences in si rates were arguably attributable to the nature of tvjt and pst. however, it is possible that the inclusion of both tasks in a within-participant design may prompt different behaviors contingent on the task type. to rule out the possibility that the effects we found were artifacts of the design, we conducted a follow-up study in which task variation was manipulated between participants. figure 2: rate of si computation estimated by tvjt and pst in experiment 1. the y-axis shows the percentage of deriving si for a given scale (“ad hoc” vs “or-and” vs “some-all”) in each task (tvjt vs pst), with zero meaning zero percent and one meaning 100 percent. confidence intervals were computed using bootstrapping methods. 3. experiment 2. in experiment 1, we observed varying estimated implicature rate contingent on the task (tvjt vs. pst) and the lexical scale (“some-all”, “or-and”, and ad hoc). however, it is possible that participants’ different responses to different tasks were an artifact of the withinsubjects design of the study. participants saw truth judgment questions and picture selection questions in a random order so they may have decided to treat them differently. in this second study, we ran the tasks between-subjects to address this possible confound. 3.1. methods. experiment 2 adopted the same set of stimuli used in experiment 1, but the tvjt items and the pst items were a between-subject manipulation rather than a within-subject manipulation. 50 participants were recruited from the online crowdsourcing platform prolific and were randomly assigned to perform either the tvjt task or pst task. of the 50 participants recruited, two were excluded for low accuracy on the attention checking control items, leaving 48 participants in the final analysis (24 in tvjt group and 24 in pst group). we combined the truth value judgment task data with the picture selection task data and adopted a similar set of analyses as in experiment 1, with minor changes to the random effect structure necessitated by the between-subjects design. since the task difference was manipulated between participants, the random effect structure of the model only included by-subject random intercepts and scaleby-subject slopes. proceedings of elm 3: 308-318, 2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi: experimental paradigms on scalar implicature estimation. 313 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3.2. results. all the main effects and the interaction we observed in experiment 1 were replicated in experiment 2 (figure 3). in the pst task, the “some-all” trials (beta = 28.07, ci = [14.00, 56.33]) and “ad hoc” trials (beta = 26.88, ci = [12.22, 56.56]) received noticeably more pragmatic interpretation than the baseline “or-and” trials. while no statistically meaningful difference in si rate was observed between tvjt and pst in the “or-and” trials (beta = 1.85, ci = [-5.47, 9.92]), for the other two scales, participants showed noticeably less si computation in tvjt compared with pst (beta = -28.83, ci = [-61.96, -10.55] for the “some-all” trials; beta = 52.16, ci = [-112.47, -23.12] for the “ad hoc” trials). moreover, there was no effect of item iteration on the probability of computing sis for both tvjt (beta = -0.47, ci = [-2.03, 0.63]) and pst (beta = 1.54, ci = [-0.88, 5.98]), the same as what we found in experiment 1. the coefficient alpha for the 18 tvjt and 18 pst critical items was 0.91 and 0.86, respectively. figure 3: rate of si computation estimated by tvjt and pst in experiment 2. the y-axis shows the percentage of deriving si for a given scale (“ad hoc” vs “or-and” vs “some-all”) in each task (tvjt vs pst), with zero meaning zero percent and one meaning 100 percent. confidence intervals were computed using bootstrapping methods. 4. experiment 3. we tested participants on the hct and compared our results with the findings of the tvjt and pst in experiment 2. 4.1. methods. we recruited 50 participants from the online crowdsourcing platform prolific. those participants were instructed to perform a hct in a single qualtrics survey. in each trial, participants were presented with two cards and a sentence potentially describing the content of one card. among the two cards, only one of them had contents visible to the participants, while the content of the other card was covered. participants were instructed to choose the card that best matched the given sentence. in the critical trials (figure 1c), the sentence was logically compatible with the visible card, but the implicature of the sentence did not match the visible card, and thus, participants’ preference to the visible card was interpreted as the absence of scalar implicature computation, but their preference to the hidden card was coded as scalar implicature computation. the stimuli used in this experiment were adopted from the same inventory of stimuli for the pst in the previous experiments with an important modification: one card in the stimuli was replaced by the “better picture” card. for the critical trials, the “better picture” card proceedings of elm 3: 308-318, 2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi: experimental paradigms on scalar implicature estimation. 314 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ always replaced the card that matched the implicature of the sentence, while for the control conditions, the “better picture” card randomly replaced one of the two cards in the trial. an example of the experimental manipulation for the hct is shown in table 1. 4.2. analysis. of the 50 participants recruited, 6 were excluded for low accuracy (accuracy rate less than 90%) on the attention checking control items, leaving 44 participants in the final analysis. since each critical item appeared twice in the experiment (following the same design as experiment 1 and 2), we constructed a bayesian logistic generalized linear model to check if the scalar implicature rate differed depending on whether the participants encountered the item for the first or the second time and calculated the composite reliability of the critical items in the hct. then we combined the data of experiment 3 with the data of experiment 2, treating the tvjt, pst, and hct as a between-subject manipulation. we constructed a bayesian logistic regression model to explore how task variation influences scalar implicature computation. the probability of computing scalar implicatures was modeled as a function of task type (tvjt vs. pst vs. hct), scale (“some-all” vs. “or-and” vs. “ad hoc”) and their interactions. the model also included by-subject random intercepts and scale-by-subject slopes. the predictors were dummy coded with the “or-and” trials of the pst as the baseline for comparison. we also plotted participants’ responses in the literal non-implicature trials across task and scale types. 4.3. results. for the experimental items, we again observed an interaction between lexical scales and task type that affected the estimated scalar implicature rate (figure 4). when the “orand” condition in the pst was set as the baseline for comparison, there was neither a statistical difference between the baseline and the “or-and” condition in the hct (beta = 0.05, ci = [-6.12, 6.37]), nor between the baseline and the same scalar items in the tvjt (beta = 2.5, ci = [-4.17, 9.69]). for the pst, both the “ad-hoc” (beta = 18.01, ci = [9.8, 30.16]) and “some-all” items (beta = 32.98, ci = [19.01, 56.33]) elicited more scalar implicature computation than the baseline “or-and” trials, however, the estimated scalar implicature rate dropped significantly for the “adhoc” (beta = -22.47, ci = [-41.7, -10.31] ) and “some-all” items (beta = -25.97, ci = [-48.8, 11.91]) in the hct and in the tvjt as well. same as the previous experiments, we did not find an effect of trial iteration on the estimated scalar implicature rate (beta = -0.24, ci = [-3.66, 3.14]). the coefficient alpha for the 18 hct critical items was 0.91. proceedings of elm 3: 308-318, 2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi: experimental paradigms on scalar implicature estimation. 315 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 4: rate of si computation estimated by hct, pst and tvjt in experiment 2 and 3. the y-axis shows the percentage of deriving si for a given scale (“ad hoc” vs “or-and” vs “some-all”) in each task (hct vs pst vs tvjt), with zero meaning zero percent and one meaning 100 percent. confidence intervals were computed using bootstrapping methods. 5. discussion. though the tvjt, pst and hct are supposed to measure the rate at which scalar implicatures are computed, the estimates from different tasks varied noticeably even for the trials that were parallel across the three tasks. the tvjt and pst were compared in a within-subject manipulation and the findings were later replicated in a between-subject manipulation. in general, the estimated si rate was much higher in the pst than in the tvjt. this was the case for both the “ad hoc” and the “some-all” scale. when reading expressions like “some of the animals are cats” or “the card has a cat”, participants clearly derived scalar implicatures as shown in their strong preference for pragmatically felicitous cards in the pst, e.g. a card with three cats and three elephants or a card with only a cat. however, they still judged the pragmatic infelicitous interpretations (e.g. a card with six cats for the description “some of the animals are cats”) as “true” for the majority of the tvjt trials regardless of the scale. this pattern supported katsos and bishop’s claim that participants are more tolerant to pragmatic infelicity than logical violations, and thus reluctant to judge under-informative statements as “false”. the form of the hct appeared to mimic the pst in that there were two pictures provided and the only difference was that a visible card in the picture selection task was replaced by the “better picture” card in the hct; however, the estimated si rate in the hct was noticeably lower than in the pst for the “ad hoc” and the “some-all” scale. for example, while participants in the pst were more likely to select the card with only a cat rather than the card with both a cat and an elephant given the prompt “the card has a cat”, they nevertheless preferred the card with a cat and an elephant when this card was paired with a “better picture” card in the hct. the patterns observed in the hct thus mimic the patterns in the tvjt more than those in the pst. while “implicature rate” as a measurement construct was reliable in repeated measurements of participants within each task, it showed high variability and therefore, low reliability as a construct between the three tasks that we investigated in this study. these results suggest that either different tasks are measuring different constructs and “implicature rate” is not measuring the same thing across tasks, or that the pragmatics of each task is affecting the “implicature rate” in systematic ways that results in considerable variability across tasks. most importantly, such variability in “implicature rate” between tasks suggests that comparing “implicature rate” between populations (e.g. children vs. adults) should also be approached with caution, since even within the same task, “implicature rate” may measure different things between populations or different populations may approach the pragmatics of each task differently. future studies should also consider task variation and reliability across different populations and discover the sources of variability in experimental measurements of scalar inferences. references barr, dale j., roger levy, christoph scheepers, and harry j. tily. 2013. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language. 68(3): 255-278. bott, lewis and emmanuel chemla. 2016. shared and distinct mechanisms in deriving linguistic enrichment. journal of memory and language, 91: 117–140. proceedings of elm 3: 308-318, 2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi: experimental paradigms on scalar implicature estimation. 316 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ bott, lewis and ira noveck. 2004. some utterances are underinformative: the onset and time course of scalar inferences. journal of memory and language, 51(3): 437–457. bürkner, paul-christian. 2017. brms: an r package for bayesian multilevel models using stan. journal of statistical software, 80(1): 1–28. chemla, emmanuel and benjamin spector. 2011. experimental evidence for embedded scalar implicatures. journal of semantics, 28(3): 359–400. chierchia, gennaro. 2013. logic in grammar: polarity, free choice, and intervention. oxford: oxford university press. chierchia, gennaro, danny fox, and benjamin spector. 2012. scalar implicature as a grammatical phenomenon. in p. portner, c. maienborn, and k. von heusinger (eds.), an international handbook of natural language meaning, vol. 3, 2297–2332. berlin: mouton de gruyter. crain, stephen and rosalind thornton. 2000. investigations in universal grammar: a guide to experiments on the acquisition of syntax and semantics. cambridge: mit press. cronbach, lee j. 1951. coefficient alpha and the internal structure of tests. psychometrika, 16: 297–334. http://dx.doi.org/10.1007/bf02310555. de neys, wim and walter schaeken. 2007. when people are more logical under cognitive load. experimental psychology, 54(2): 128–133. https://doi.org/10.1027/1618-3169.54.2.128. degen, judith and michael k. tanenhaus. 2015. processing scalar implicature: a constraintbased approach. cognitive science, 39(4): 667–710. https://doi.org/10.1111/cogs.12171. degen, judith and michael k. tanenhaus. 2019. constraint-based pragmatic processing. in c. cummins and n. katsos (eds.), the oxford handbook of experimental semantics and pragmatics, 485–504. oxford: oxford university press. doran, ryan, gregory ward, meredith larson, yaron mcnabb, and rachel e. baker. 2012. a novel experimental paradigm for distinguishing between what is said and what is implicated. language, 88(1): 124–154. feng, shuo and jacee cho. 2019. asymmetries between direct and indirect scalar implicatures in second language acquisition. frontiers in psychology, 10: 877. gamut, l. t. f. 1991. logic, language, and meaning. volume i: introduction to logic. chicago: university of chicago press. gazdar, gerald. 1979. pragmatics: implicature, presupposition and logical form. new york: academic press. gerken, louann and mary e. shady. 1998. the picture selection task. in d. mcdaniel, c. mckee, and h. s. cairns (eds.), methods for assessing children’s syntax, 157–173. cambridge: mit press. geurts, bart and nausicaa pouscoulous. 2008. no scalar inferences under embedding. in p. egre and g. magri (eds.), presuppositions and implicatures, 1–22. mit working papers in linguistics. geurts, bart and nausicaa pouscoulous. 2009. embedded implicatures? semantics and pragmatics, 2: 1–34. gordon, peter. 1998. the truth-value judgment task. in d. mcdaniel, c. mckee, and h. smith (eds.), methods for assessing children’s syntax, 211–228. cambridge: mit press. grice, herbert paul. 1975. logic and conversation. in p. cole and j. l. morgan (eds.), syntax and semantics 3: speech acts, 41–58. new york: academic press. grice, herbert paul. 1978. further notes on logic and conversation. in p. cole (ed.), pragmatics, 113–128. leiden: brill. proceedings of elm 3: 308-318, 2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi: experimental paradigms on scalar implicature estimation. 317 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ grodner, daniel j., nadine m. klein, kathleen m. carbary, and michael k. tanenhaus. 2010. “some,” and possibly all, scalar inferences are not delayed: evidence for immediate pragmatic enrichment. cognition, 116(1): 42–55. hirschberg, julia. 1985. a theory of scalar implicature. ph.d. dissertation, university of pennsylvania, philadelphia, pa. http://repository.upenn.edu/dissertations/aai8603648. horn, laurence. 1972. on the semantic properties of logical operators in english. ph.d. dissertation, university of california, los angeles, ca. horowitz, anna c., rebecca m. schneider, and michael c. frank. 2018. the trouble with quantifiers: exploring children's deficits in scalar implicature. child development, 89(6): e572– e593. huang, yi ting and jesse snedeker. 2009. online interpretation of scalar quantifiers: insight into the semantics-pragmatics interface. cognitive psychology, 58(3): 376–415. https://doi.org/10.1016/j.cogpsych.2008.09.001. huang, yi ting and jesse snedeker. 2011. logic and conversation revisited: evidence for a division between semantic and pragmatic content in real-time language comprehension. language and cognitive processes, 26(8): 1161–1172. katsos, napoleon and dorothy v. m. bishop. 2011. pragmatic tolerance: implications for the acquisition of informativeness and implicature. cognition, 120: 67–81. https://doi.org/10.1016/j.cognition.2011.02.015. kuder, g. frederic and marion w. richardson. 1937. the theory of the estimation of test reliability. psychometrika, 2: 151–160. https://doi.org/10.1007/bf02288391. levinson, stephen c. 1983. pragmatics. cambridge: cambridge university press. levinson, stephen c. 2000. presumptive meanings: the theory of generalized conversational implicature. mit press. noveck, ira a. and alicia posada. 2003. characterizing the time course of an implicature: an evoked potentials study. brain and language, 85: 203–210. papafragou, anna and julien musolino. 2003. scalar implicature: experiments at the semanticspragmatics interface. cognition, 86: 253–282. https://doi.org/10.1016/s00100277(02)00179-8. potts, christopher, daniel lassiter, roger levy, and michael c. frank. 2015. embedded implicatures as pragmatic inferences under compositional lexical uncertainty. journal of semantics, 33(4): 755–802. https://doi.org/10.1093/jos/ffv012. rees, anna and lewis bott. 2018. the role of alternative salience in the derivation of scalar implicatures. cognition, 176: 1–14. https://doi.org/10.1016/j.cognition.2018.02.024. rees, anna, emily carter, and lewis bott. 2023. priming scalar and ad hoc enrichment in children. cognition, 239: 105572. https://doi.org/10.1016/j.cognition.2023.105572. slabakova, roumyana. 2010. scalar implicatures in second language acquisition. lingua, 120(10): 2444–2462. sperber, dan and deirdre wilson. 1995. relevance: communication and cognition. oxford: blackwell. van tiel, bob, eva van miltenburg, natalia zevakhina, and bart geurts. 2016. scalar diversity. journal of semantics, 33(1): 137–175. proceedings of elm 3: 308-318, 2025 zhuang qiu, casey d. felton, zachary n. houghton, and masoud jasbi: experimental paradigms on scalar implicature estimation. 318 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ non-implicature sources of exclusivity in linguistic disjunction casey felton & masoud jasbi * abstract. disjunction in natural language alternates between an inclusive reading (a or b or both) and an exclusive reading (a or b but not both). traditional accounts of this ambiguity focus on scalar implicature as the source of disjunction exclusivity, a process whereby gricean reasoning over horn scales strengthens the baseline inclusive reading to an implied exclusive reading (grice, 1978; horn, 1972; gazdar, 1980). despite nearly all theories acknowledging that other factors likely play a role in the generation of exclusivity implications, non-implicature factors have received comparatively little attention. across four experiments we tested two such non implicature factors, prior compatibility and syntactic category, finding that both play a role in speaker interpretations of disjunctive sentences. additionally, by drawing our stimuli in the first two experiments from the prior literature, we found evidence that previous research on disjunction, while accurately identifying the key role of scalar implicatures, may be overestimating the effect size thereof due to a failure to control for non-implicature factors. keywords. experimental pragmatics; prior probability, scalar implicature, disjunction 1. introduction. when a listener encounters a natural language disjunction they must often determine whether the speaker intended an inclusive (a or b or both) meaning or an exclusive (a or b but not both) meaning. the most studied mechanism explaining the derivation of these implications is scalar implicature; a gricean process wherein reasoning about the alternative words a speaker could have chosen, but did not chose, allows a listener to infer the logically stronger exclusive reading in many instances of natural language disjunction. that said, numerous other factors likely contribute to the interpretation of disjunction. one such factor that intuitively must contribute is prior compatibility, by which we mean a speaker’s preconceived notions of the likelihood that two disjuncts would be true together. cases exist wherein two disjuncts that are not logically incompatible are nonetheless extremely unlikely to co-occur, such as the sentence john is singing or screaming, which we sampled from the literature. while john may be doing both, it seems far more likely that someone would not be both singing and screaming at the same time. this prior incompatibility potentially limits the amount of extra exclusivity that can be introduced by scalar reasoning and other non-implicature factors. extreme examples, such as those noted by geurts (2006), may even lead to instances of exclusivity implications sans any scalar reasoning, e.g. in a sentence like the ball is on the table or on the floor, a speaker can reasonably derive an exclusive reading from their knowledge that objects can only be in a single location at a time. another factor worth considering is the syntactic categories of the disjuncts themselves. authors: casey felton, university of california, davis (cdfelton@ucdavis.edu) & masoud jasbi, university * of california, davis (jasbi@ucdavis.edu). proceedings of elm 3: 163-175, 2025 c©2025 casey felton and masoud jasbi published by the lsa with permission of the author(s) under a cc by license. 163 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ for many disjunctions, the same semantics can be expressed with multiple syntactic frames, as in the following examples: (1) a. john likes coffee or tea. b. john likes coffee or john likes tea. (jasbi, 2018) to some speakers, sentences formed by coordinating clauses lead to more exclusive interpretations than sentences formed by coordinating nps. due to the flexibility of english coordination, it is typically possible to generate object position disjuncts through the coordination of clauses, nps, and vps, but to date no experimental work has been done investigating whether speakers differ in their interpretation of these semantically identical but syntactically distinct sentences. in this paper we present four experiments that test the role of two non-implicature factors on the generation of exclusivity implications. in experiments 1 and 2 the role of prior compatibility, as well as the degree to which the literature on exclusivity implications in natural language disjunction controls for it, was assessed. in experiments 3 and 4 the role of the varying syntactic categories of disjuncts coordinated by “or” was assessed while controlling for disjunct length. 2. prior compatibility. experiments 1 and 2 tested the extent to which prior compatibility of disjuncts may have contributed to the exclusivity of examples and stimuli used in the literature on scalar implicatures. while the literature on scalar implicatures has acknowledged the role prior compatibility of disjuncts may play in generating exclusivity implications, there has been no study that measures it systematically. in experiment 1, we sampled stimuli from previously published theoretical and experimental work on exclusivity implications and asked participants to rate disjuncts independently and without the presence of the connective “or” with respect to their compatibility. if prior compatibility of disjuncts had no contribution to exclusivity in these examples, we expected participants to rate the disjuncts separately as compatible. in experiment 2, we asked a different group of participants to use the same scale to rate the disjunctive examples, this time with the word “or” present in the sentences. by the design of the two studies, the stimuli differed mainly by the addition of the word “or”, providing a measure of the exact amount of exclusivity, if any, that is attributable to scalar reasoning over and above the effect of prior compatibility measured in experiment 1. we selected a between subjects design due to concerns that exposure to disjuncts without “or” in trials measuring compatibility might bias participants when exposed to the same disjuncts in the exclusivity rating task, either by causing them to match their responses or by artificially emphasizing or and thus inflating scalar implicature rate. 2.1. shared stimuli. we sampled 47 examples from prior literature using the following criteria. first, a disjunction had to be about something that it is reasonable for speakers to have world knowledge about. for instance, the following sentence from quelhas & johnson-laird (2017) was excluded: teresa is writing a poem or joaquim is painting a picture. this was not suitable for our experiment because there is no reason to expect speakers to have any prior knowledge about the characters teresa and joaquim. similar exclusions were made for items with disjuncts that are unrelated to one another, such as the bird is in the nest or the shoe is on the foot (paris, 1973). second, disjunctions could not be interrogatives, as the design of the study was based proceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 164 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ around declarative or imperative sentences. finally, the paper the disjunction occurred in had to discuss the alternation between inclusive and exclusive readings of or, broadly construed. we did not exclude examples of embedded disjunction or those with varying entailment environments which potentially introduced more by-item variability to the results of experiment 2 (see our web resources at https://doi.org/10.17605/osf.io/s2j9d for a full list of stimuli and their sources). the final set of stimuli, which were shared across the first two experiments, consisted of 47 disjunctions taken from 20 published papers. most items were originally presented in english, but those taken from katsos, breheny & williams (2005) were translated from greek. minor changes to the original sentences were sometimes necessary to integrate these items into our experimental designs. alongside our experimental items, six control items were also created, three with disjuncts that cannot co-occur and three with disjuncts that must co-occur. these control items served as attention checks because their compatibility or incompatibility was unambiguous. this meant a total of 53 items, 47 experimental items and 6 controls (see (2) for examples). (2) a. example experimental item: john brought pizza or pasta to the party. b. example control item (incompatible): the defendant is innocent of all charges or guilty of all charges. c. example control item (highly compatible): lauren is alive or not dead. all items that were hypothesized in their source articles to be more likely to elicit an exclusive or inclusive interpretation (widely construed) were annotated as either hypothesized exclusive or hypothesized inclusive. twenty-five items were hypothesized exclusive, fifteen were hypothesized inclusive, six were controls, and seven did not have a predicted reading in their source, but were nonetheless judged to be quality stimuli and thus were included. 2.2. experiment 1 methods. in experiment 1 we conducted an online behavioral study to collect data on subjects’ prior beliefs regarding the compatibility of the disjuncts in the stimuli. in this experiment participants did not see any instance of linguistic disjunction; they were presented only with the disjuncts themselves without or. participants. we recruited 57 participants from the prolific participant pool. all participants were over the age of 18, and self reported as being both first language english speakers and american nationals. of the 57 initially recruited, six declined to participate and one timed out automatically, leaving 50 participants who submitted responses and received $3.75 in compensation. forty-two responses were included in the analysis after excluding those who answered too many attention checking control trials incorrectly. stimuli. for the experiment 1 stimuli, we removed or from the examples of disjunction we had collected from the literature, making sure that number agreement was adjusted as needed to maintain grammaticality. next, we generated a question for each item that asked how likely it was that both of the two disjuncts were true together, that is, asking how compatible the items were. the example experimental item in (3) below explains our process. for example, the disproceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 165 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ junction john brought pizza or pasta to the party would be broken down into the two disjuncts (3a) and (3b) and paired with the question (3c). (3) a. disjunct a: john brought pizza to the party. b. disjunct b: john brought pasta to the party. c. question: how likely is it that someone brought both pizza and pasta to the party? procedure. experiment 1 consisted of 53 randomly ordered trials, each consisting of 2 paired disjunct sentences, the corresponding likelihood question, and a slider bar to respond to the question. after consenting to participate in an online behavioral experiment, participants read instructions and completed 6 practice trials with feedback before beginning the real experiment. the experiment was self paced and took 12 minutes on average to complete. in each trial participants were instructed to read the two disjunct sentences and then respond to the question using the slider. the slider ranged from 0% (it is impossible for the two disjuncts to co-occur) to 100% (the two disjuncts must co-occur). to reduce rushed responses, the slider had to be moved in order for a participant to proceed to the next trial. the feedback given on practice trials was designed to make sure participants had a good grasp of the scale and were aware that they could rate trials exactly 50% by moving the slider back to its starting position. 2.3. experiment 2 methods. in experiment 2 we aimed to collect exclusivity judgements for the same 47 disjuncts tested in experiment 1. unlike experiment 1, here we presented the full examples of disjunction with the connective or. participants. we recruited 52 participants from the prolific participant pool. all participants were over the age of 18, and self reported as being both native speakers of english and american nationals. of the 52 initially recruited, one declined to participate and one timed out due to inactivity, leaving 50 participants who submitted complete responses and received $3.75 in compensation. forty-one responses were included in the analysis after excluding those who answered too many attention checking control trials incorrectly. stimuli. the stimuli for experiment 2 were the examples of disjunction collected from the literature on the semantics and pragmatics of disjunction. four examples had the word either in addition to the disjunction word or. since the role of either was not in the scope of this study, it was removed from these four stimuli. unlike in experiment 1 where number agreement was adjusted to maintain grammaticality, here number agreement was left as it appeared in the literature. like experiment 1, each item was paired with a question, this time designed to assess the overall exclusivity of a linguistic disjunction. for example: (4) a. full disjunction: john brought pizza or pasta to the party. b. question: based on the sentence, how possible is it that john brought both pizza and pasta to the party? in order to draw participants’ attention to the linguistic disjunction, all questions began with “based on the sentence, . . . ” and the word “possible” was swapped for “likely.” these changes proceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 166 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ in the wording of the question aimed to help participants focus on the linguistic assessment of disjunction rather than their prior beliefs about the compatibility of the events described. procedure. experiment two consisted of 53 randomly ordered trials, each consisting of a full disjunction sentence, the corresponding question, and a slider bar to respond to the question. after consenting to participate in an online behavioral experiment, participants read instructions and completed 3 practice trials with feedback before beginning the real experiment. the experiment was self paced and took 10 minutes on average to complete. in each trial participants were instructed to imagine that someone said the disjunction sentence and then respond to the question using the slider. a slight change in instructions to “imagine someone said the sentence” was made to further draw participants’ attention to the linguistic nature of the task. the task was otherwise identical to that in experiment 1. 2.4. experiment 1 results. figure 1 shows the compatibility ratings provided by the participants in experiment 1 on the y axis, and is sorted by item on the x axis. it is clearly visible by eye that the items predicted to generate exclusivity implicatures in the literature are grouped around the lower end of the ratings, with the opposite being the case for those predicted to be biased towards an inclusive interpretation (note that blue items are grouped towards the right half of figure 1 while green items are grouped towards the left). figure 1: experiment 1 ratings (lowest to highest) while both a linear regression or a beta regression are reasonable models to statistically assess whether the items with different hypothesized readings differ in mean compatibility, we selected a bayesian beta regression to allow for a more accurate modeling of the boundedness of the data. after inflating 0s and deflating 1s to 0.0001 and 0.9999 respectively, both the central tendency and spread of the responses were modeled as a function of an item’s claimed exclusivity or inclusivity in the literature and a random intercept by participant. all r-hats were 1 equal to 1, and inspection of tranq plots suggested no issues with mixing. the model showed a higher mean value for items predicted by the literature to be inclusively biased (β = 0.88, 95% credibility interval = [0.77, 1]), which confirms that the items predicted to generate exclusivity because no qualities of the raw data or theoretical considerations suggested that participants treated or ought 1 to treat ratings of 0 or 1 as qualitatively distinct from ratings of 0.01 or 0.99, we decided to eschew a 0 & 1 inflated beta regression in order to reduce model complexity. proceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 167 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ implicatures tended to be rated as less compatible while items predicted not to generate implicatures were rated as more compatible, despite no disjunction word being used and therefore no scalar reasoning being possible. this suggests that items in the literature were not being controlled for prior compatibility, and that any claims based off of them may be confounded by it. 2.5. experiment 2 results. figure 2 shows participants’ ratings of exclusivity for the linguistic items in experiment 2. while there is clear variability among the items, it is easy to see that overall the majority of items are rated lower in experiment 2 than in experiment 1 despite the only major difference between tasks being the addition of the disjunction word or. in addition, the predicted interpretations from the theoretical literature generally line up with the interpretations provided by participants naive to the linguistic theory. finally, while some items showed an increase in ratings from experiment 1 to experiment 2, this is likely attributable to between group variation in prior compatibility norms. figure 2: experiment 2 ratings (ordering identical to figure 1) by regressing the mean compatibility ratings on the mean inclusivity ratings (analogous to implication rate) we were able to directly assess the role of compatibility on exclusivity implication generation. as figure 3 shows, items that were rated as more compatible were subsequently rated as more inclusive; or less exclusive. traditional statistics lack straightforward ways to compare the results of the two experiments, but hierarchical bootstrapping as outlined by saravanan, berman, & sober (2020) allows an estimation of the correlation between item means in each experiment. ten-thousand bootstrapped samples were conducted, returning a mean correlation of 0.527 with a 95% ci of [0.439, 0.608]. this supports the hypothesis that prior beliefs regarding the compatibility of the two disjuncts contribute significantly to the overall interpretation of a disjunction as exclusive. it is important to note here that our stimuli come from the theoretical literature on scalar implicature, where researchers ostensibly select examples that are not a-priori exclusive. nevertheless we see that prior beliefs of compatibility do have a significant effect. it is quite possible that the actual role of prior beliefs on judgments of exclusivity is even more prominent in every-day examples. going beyond the predictions made by geurts (2006) that mutually incompatible disjuncts are interpreted exclusively sans implicature, our results suggest not only that disjuncts judged less likely to co-occur often lead to a more exclusive interpretation, but also that disjuncts judged highly likely to co-occur often lead to a more inclusive interpretation. additionally, the effect does not seem to be limited to completely muproceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 168 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ tually incompatible disjuncts, but extends to differences in prior compatibility across a range of values. even though a large portion of the variance in disjunction exclusivity was explained by variance in prior beliefs about disjunct compatibility, there remained a large amount of variance attributable to other factors, such as pragmatic scalar exclusivity implicatures. for most items, especially those hypothesized to be interpreted as exclusive by the literature, the ratings tended to be lower (that is, more exclusive) in experiment 2. these differences cannot be attributed to compatibility, and are likely attributable to scalar implicatures given that the two experiments mainly differed with respect to their inclusion of linguistic disjunction, though we cannot rule out other non-implicatures factors as contributors. figure 3: correlation of experiment 1 and 2 ratings 3. syntactic category. 3.1. experiment 3 methods. in experiment 3 we aimed to assess the effects of syntactic category and disjunct length while controlling for prior compatibility. additionally, we included an exploratory condition for whether the use of a pronoun rather than repeating the proper name of the subject impacted the interpretation of coordinated clauses, e.g. john sings or he screams, versus john sings or john screams. experiment 3 uses a similar procedure to experiment 2, but with stimuli that control for disjunct length and syntactic category. proceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 169 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ participants. we initially recruited 64 participants from the prolific participant pool, but later recruited an additional 11 to increase power (since we use a bayesian analysis this does not violate the assumptions of our statistics). all participants were over the age of 18, and self reported as being both native speakers of english and american nationals. of the 75 total participants recruited, 3 declined to participate and 2 timed out due to inactivity, leaving 70 participants who submitted complete responses and received $3.75 in compensation. 52 responses were included in the analysis after excluding those who answered too many attention checks and control trials incorrectly. stimuli. the stimuli for experiment 3 were not pulled from previous literature, but instead were created for the explicit purpose of controlling for syntactic category and disjunct length. 32 sentence frames were created that can appear in any of the 4 conditions shown in the table below, but which varied in np length from 1 to 8 words. frames were sorted into 4 groups to ensure that each participant only saw each frame in a single permutation. an additional 18 filler items were created that used “and” or “but not” as connectives. these fillers also served as attention checks due to their objective answers. example experimental stimuli are shown in (5). (5) a. clauses + proper name: john likes coffee or john likes tea. b. clauses + pronoun: john likes coffee or he likes tea. c. verb phrases: john likes coffee or likes tea. d. noun phrases: john likes coffee or tea. procedure. experiment 3 consisted of 50 randomly ordered trials, each consisting of a full disjunction sentence (32 items) or filler (18 items), the corresponding question, and a slider bar to respond to the question. after consenting to participate in an online behavioral experiment, participants read instructions and completed 3 practice trials with feedback before beginning the real experiment. the experiment was self paced and took 10 minutes on average to complete. in each trial participants were instructed to imagine that someone said the stimuli sentence and then respond to the question using the slider. the wording of the prompts were changed from experiment 2 to use “likely” instead of “possibly” to avoid any potential confusion stemming from a modal. as in the previous two studies, the slider had to be moved in or2 der for a participant to proceed to the next trial. 3.2. experiment 3 results. the results of experiment 3 were highly noisy, with most visual summaries providing no additional clarity. analysis of mean ratings in table 1 showed a slight trend in the hypothesized direction, with vps and nps being rated as slightly more inclusive on average than coordinated clauses, regardless of pronoun versus proper name use. no trends were observable across items that differed in disjunct length. the large amount of noise in the data, as well as the small effect size, made extracting clear conclusions difficult, but some of the effects visible in the means summary can be supported by modeling. this change was deemed appropriate because the new items are perfectly controlled for prior compatibility 2 because they are semantically identical. likelihood can thus stand in as a measure of exclusive interpretation, unlike in experiment 2 where we wanted to specifically emphasize the disjunction while de-emphasizing compatibility. proceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 170 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ table 1: experiment 3 mean ratings by syntactic category the modeling challenges in experiment 3 run parallel to those in experiments 1 and 2 due to the similarly unusual distribution of the results. in comparison to the previous experiments there is far less one inflation, but instead a much higher rate of zero inflation. inspection of the data revealed that this zero inflation seems to be the result of some participants having a large bias towards “0%” responses with little variance in their data, nearly always responding that there was no chance both disjuncts were true. due to this, we elected to use a zero inflated beta regression that models the “0%” responses separately from the non-zero responses, since many of them seemed to be stemming from a different process than the non-zero responses. additionally, the differences in the shape of the distribution of data for each condition are minimal, suggesting that it is likely unnecessary to model the precision of the beta distribution as a function of our variables. these considerations led us to the following model which aimed to test the minimal aim of the experiment: are coordinated clauses interpreted more exclusively than coordinated vps and nps, and is this effect moderated by constituent length. µ ~ constituent_length*syntactic_category + (1|item) + (1|participant) φ ~ 1 + (1|item) + (1|participant) zi ~ 1 + (1|item) + (1|participant) in the model, syntactic category was collapsed to compare just the two clause conditions (proper noun and pronoun) to the two non-clause conditions (vps and nps), and constituent length was grand mean centered. four chains of 4000 samples with 1000 warm up samples were used; all r-hats approached 1, and inspection of tranq plots suggested no issues with mixing. the model output suggests a small positive effect of constituent type (β = 0.09, 95% credibility interval = [0, 0.19]). inspection of the posterior samples show that 96.9% of the samples estimated an effect of non-clausal syntactic category that was positive. no meaningful effect of any other predictor or interaction was observed, reinforcing what was visible by eye: constituent length both plays no observable role in exclusivity implications and does not confound the effect of syntactic category. that said, the observed effect of syntactic category is both quite small in magnitude and has a credibility interval that comes close to crossing zero, which motivated us to conduct experiment 4 in order to determine whether the effect of syntax on implication generation is real but small in magnitude or spurious. 3.3. experiment 4 methods. the results of experiment 3 left in question whether there is a real effect of syntactic category with a small effect size or no effect of syntactic category at all. experiment 4 directly tests this question in a modified picture selection task (gerken & shady, 1998) that striped away all potential confounds and directly assessed whether disjunctions with mean score s.e. clauses (proper noun) 31.15 3.862320 clauses (pronoun) 29.06 3.756284 verb phrases 32.13 3.839423 noun phrases 34.02 4.119652 proceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 171 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ coordinated nps are interpreted more inclusively than semantically matched disjunctions with coordinated clauses. participants. we recruited 51 participants from the prolific participant pool. all participants were over the age of 18, and self reported as being both native speakers of english and american nationals. of the 51 initially recruited, 1 declined to participate and no paritcipants timed out, leaving 50 participants who submitted complete responses and received $3.75 in compensation. 49 responses were included in the analysis after excluding 1 participant who answered too many attention checks and control trials incorrectly. stimuli. the stimuli for experiment 4 were highly simplified compared to those in experiment 3, and were based on those used in qiu et al. (2023). a total of twenty four sentences were created made up of 12 experimental disjunction sentences (half clauses, half nps) and 12 control items. the control sentences had objectively correct answers, allowing them to serve as attention checks as well as distractors. each sentence was paired with 4 image cards, like those in figure 4. two of the cards contained a single animal mentioned in the sentence, one of the cards contained both of the two animals mentioned in the sentence, and the final card was always blank. participants were instructed to select one or more cards to reflect their interpretation of the stimuli sentence. this design allows participants to respond in a way that expresses the uncertainty that is central to the meaning of disjunction, unlike traditional picture selection tasks which do not. figure 4: experiment 4 trial example procedure. experiment 4 consisted of 36 randomly ordered trials, each consisting of a disjunction item (24 items, each trial presented twice) or a control item (12 items) and the 4 images paired with that item. after consenting to participate in an online behavioral experiment, participants read instructions and completed 3 practice trials before beginning the real experiment. the practice items did not employ “or” in order to avoid bias, but did include examples that required the selection of multiple cards in order to ensure that participants were aware that they could select more than one card. non-practice items were presented in a random order, and the linear order of the image cards were scrambled in each trial. in each trial, participants read a proceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 172 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ stimuli sentence and then selected one or more cards based on their interpretation of it. the experiment was self paced and took less than 5 minutes to complete on average. 3.4. experiment 4 results. after excluding participants with low accuracy on control items, the data was filtered to include just trials involving and and or for comparison. the responses were scored categorically into conjunction (selecting just the card with 2 animals), exclusive disjunction (selecting just the two cards with a single animal each), inclusive disjunction (selecting all the cards except the blank one) and other (any other response). a complete summary of the responses is shown in figure 5, where it is clear by eye that the overwhelming majority of and trials were responded to correctly (participants provided conjunctive responses). the or trials were split between inclusive and exclusive disjunctive responses, with slightly more inclusive responses in the coordinated np condition visible by eye. figure 5: summary of e4 responses to and & or items a bayesian logistic mixed effects regression was fit to test the effect of syntactic category on participant responses to or items. responses were modeled as a function of the syntactic category of their disjuncts, alongside a random intercept of both item and participant and a random slope of syntactic category by participant. four chains of 4000 samples with 1000 warm up samples were used; all r-hats approached 1, and inspection of tranq plots suggested no issues with mixing. the model output suggests a non-zero effect of syntactic category (β = -2.06, 95% credibility interval = [-4.14, -0.20]), meaning that the exclusivity rate was lower for items with coordinated nps than for items with coordinated clauses, confirming our initial hypotheses. that said, the effect size is small, the model predicting only an 11% change in exclusivity between item types. taken together with the results of experiment 3, our data suggests that there is a real, but small effect of syntactic category on exclusivity implication generation. 4. conclusions. across 4 experiments we investigated 2 non-implicature sources of exclusivity. in experiment 1, we used examples from the literature on scalar implicatures to measure the proceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 173 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ degree to which prior compatibility contributes to judgments of exclusivity. the results showed that sentences that were judged to be exclusive tended to have disjuncts that were judged to be less compatible and that those judged inclusive tended to have disjuncts that were judged to be more compatible. this finding suggests that informal judgments of exclusivity in the theoretical literature may risk being confounded, unless properly controlled in appropriate minimal-pair contrasts. the results of experiment 2 confirmed that prior compatibility plays a role in exclusivity implication, explaining a large portion of the overall variance in participant ratings of full disjunctions. that said, the results also clearly demonstrated that the mere presence of the disjuncts with the word "or" in a disjunction introduces exclusivity implications over and beyond judgments of prior compatibility; an effect most likely attributable to scalar implicatures generated by the usage of disjunction word "or". next, experiment 3 tested the role of syntactic category and disjunct length, finding no evidence for an effect of disjunct length, but marginal evidence that coordinated nps may be interpreted more inclusively than coordinated vps or clauses. finally, in experiment 4 we re-tested the effect of syntactic category with a new task designed to reduce noise and detect a lower magnitude effect. the results suggested a consistent, but small, effect of syntactic category, with coordinated clauses being interpreted more exclusively than coordinated nps. further research should aim to clarify the theoretical model implied by our data, as well as investigate potential interactions between the sources of exclusivity implications in natural language. references breheny, richard, napoleon katsos & john williams. 2005. interaction of structural and contextual constraints during the on-line generation of scalar inferences. proceedings of the annual meeting of the cognitive science society, 27. https://escholarship.org/uc/item/73c75114 gazdar, gerald. 1980. pragmatics and logical form. journal of pragmatics 4(1):1-13. https:// doi.org/10.1016/0378-2166(80)90014-4 geurts, bart. 2006. exclusive disjunction without implicature. nijmegen: university of nij megen unpublished ms. note. geurts, bart. 2010. quantity implicatures. cambridge: cambridge university press. gerken, louann, & shady, michelle e. 1998. the picture selection task. in dana mcdaniel, cecil mckee, & helen. s. cairns (eds.), methods for assessing children's syntax. cambridge: mit press. grice, herbert p. 1978. further notes on logic and conversation. in pragmatics (pp. 113-127). leiden: brill. horn, laurence r. 1972. on the semantic properties of logical operators in english. los angeles, ca: university of california dissertation. jasbi, masoud. 2018. learning disjunction. stanford, ca: stanford university dissertation. paris, scott g. 1973. comprehension of language connectives and propositional logical relationships. journal of experimental child psychology, 16. 278-291. https:// doi.org/10.1016/0022-0965(73)90167-7 qiu, zhuang, casey d. felton, zachary n. houghton & masoud jasbi. 2023. task variation in scalar implicature computation. poster presented at the annual meeting proceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 174 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ of the cognitive science society (cogsci). quelhas, ana cristina & phillip n. johnson-laird. 2017. the modulation of disjunctive asser tions, the quarterly journal of experimental psychology, 70(4). 703-717. https://doi.org/ 10.1080/17470218.2016.1154079 saravanan, varun, gordon j. berman & samuel j. sober. 2020. application of the hierarchical bootstrap to multi-level data in neuroscience. neurons, behavior, data analysis, and theory, 3. https://nbdt.scholasticahq.com/article/13927-application of-the-hierarchical-bootstrap-to-multi-level-data-in-neuroscience. proceedings of elm 3: 163-175, 2025 casey felton and masoud jasbi: non-implicature sources of exclusivity in linguistic disjunction. 175 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ an experimental investigation of perspective alignment in gesture and speech sebastian walter & stefan hinterwimmer* abstract. hinterwimmer et al. (2021) experimentally investigated the hypothesis that perspective in gesture and speech is by default aligned, i.e., when a character’s or protagonist’s perspective is conveyed in the speech signal, this utterance is preferably aligned with a character viewpoint gesture. if an utterance expresses an observer’s perspective, by contrast, it is more likely accompanied by an observer viewpoint gesture. their results, however, showed an overall preference for character viewpoint gestures. they argued that there were pragmatic factors (e.g., informativity) at play blocking the hypothesized perspective alignment. the study reported here further investigates hinterwimmer et al.’s (2021) hypothesis by comparing two different character viewpoint gestures paired with a verbal utterance in a rating study. the results suggest that, contrary to hinterwimmer et al.’s (2021) hypothesis, multiple, potentially non-aligned perspectives can be simultaneously expressed in gesture and speech. keywords. perspective in gesture; free indirect discourse; interactions of perspective taking in gesture and speech 1. introduction. perspective plays a decisive role in the interpretation of many lexical items. the evaluative expression fantastic in (1), for example, belongs to the family of so-called perspectivedependent expressions. (1) the conference was fantastic! it is apparent that an evaluative expression as in (1) is by default interpreted from the speaker’s perspective (harris 2012). however, sometimes perspective-dependent expressions can occur in contexts where they do not depend on the perspective of the speaker of the current utterance situation: (2) a. marvin said: “the conference was fantastic!” b. marvin said that the conference was fantastic. in the direct discourse utterance (2-a) as well as the indirect discourse utterance (2-b), fantastic is interpreted from marvin’s rather than the speaker’s perspective. perspective can also be expressed in gesture (mcneill 1992). the main distinction to be drawn here is the one between character and observer viewpoint gestures. character viewpoint gestures depict an event from an internal, i.e., first-person perspective. observer viewpoint gestures, by contrast, depict events from a more external, third-person perspective. imagine that a speaker *acknowledgments: this research was conducted in the dfg-funded project visual and non-visual means of perspective taking in language which is part of the priority program 2392 (vicom). we thankfully acknowledge the financial support of the german research foundation (dfg). moreover, we would like to thank our actor lennart klappstein for the enactment of the experimental material of the study reported in this paper. moreover, we would like to thank our student assistant lennart fritzsche for helping with the video recordings. finally, we express our special gratefulness to cornelia ebert for discussing the topics covered in this paper. authors: sebastian walter, university of frankfurt (s.walter@em.uni-frankfurt.de) & stefan hinterwimmer, university of hamburg (stefan.hinterwimmer@uni-hamburg.de). proceedings of elm 3: 411-422, 2025 c©2025 sebastian walter and stefan hinterwimmer published by the lsa with permission of the author(s) under a cc by license. 411 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ wants to describe an event where someone had to run. when they want to include a gesture in that description, they could either perform a character viewpoint gesture where they depict the running with their whole body (i.e., as if they were running on one spot). alternatively, they could produce an observer viewpoint gesture where they just use their index and middle finger to represent the running person’s legs and thus depict the running event from an external perspective. research on interactions of perspective taking in gesture and speech is scarce. hinterwimmer et al. (2021) report a study investigating the hypothesis that perspective in gesture and speech is by default aligned. contrary to their hypothesis, however, the results showed an overall preference for character viewpoint gestures. they conclude that there might have been pragmatic factors blocking perspective alignment in their experimental stimuli. the study reported in this paper controls for these intervening pragmatic factors. contrary to the hypothesis, however, the results still suggest no strict preference for perspective alignment in gesture and speech. instead, multiple, potentially non-aligned perspectives can be expressed simultaneously in the two modalities. the paper is structured as follows: section 2 will provide the relevant background on perspective in speech, gesture, and previous research on interactions of the expression of perspective in the two modalities. in section 3, the study is described which investigates hinterwimmer et al.’s (2021) weakened hypothesis that in the absence of intervening pragmatic factors, perspective in gesture and speech is aligned. finally, section 4 offers a general discussion of the results and gives implications for future research. 2. background. 2.1. perspective in speech. the perspective conveyed in an utterance is normally that of the speaker. therefore, perspective-dependent expressions, such as epithets (e.g., that bastard), relational expressions (e.g., left, right, this, that), predicates of personal taste (e.g., tasty), or epistemic modals (e.g., must) are by default interpreted from the speaker’s perspective (harris & potts 2009, harris 2012). however, one can find systematic exceptions to this general tendency, namely instances of reported speech: (3) a. on her way home, mary heard a song by kendrick lamar that she liked on the radio. she thought: “i will buy his new album tomorrow.” b. on her way home, mary heard a song by kendrick lamar that she liked on the radio. she thought that she would buy his new album on the following day. c. on her way home, mary heard a song by kendrick lamar that she liked on the radio. she would buy his new album tomorrow. (hinterwimmer 2017, p. 284) in the instance of direct discourse (dd) in (3-a), all perspective-dependent expressions shift toward the reported speaker and are thus interpreted from their perspective, i.e., mary’s. this means, then, that the first-person pronoun i refers to mary, not the speaker in the current utterance situation. by contrast, the picture for indirect discourse (id) and free indirect discourse, fid, ((3-b) and (3-c), respectively) is somewhat less clear. in id (cf. (3-b)) epithets, evaluative expressions, and predicates of personal taste can in principle either be evaluated from the speaker’s perspective or from the perspective of the matrix clause subject (mary), the latter being the default option. personal pronouns (in many languages, but see proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 412 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ anand & nevins 2004), temporal, and local deictic expressions are also evaluated from the matrix subject’s perspective. for the former two, however, this preference can be easily overwritten (plank 1986, anderson 2019). by contrast, expressives and appositives are normally interpreted from the speaker’s perspective, but a shifted interpretation is also possible (harris & potts 2009). finally, the second sentence in (3-c) is an instance of fid, a way to report a protagonist’s thoughts or speech without any overt marking (e.g., hinterwimmer 2024). here, all perspectivedependent expressions are interpreted from the protagonist’s perspective with only two exceptions: pronouns and tenses (e.g., schlenker 2004). this explains why the past tense marking can cooccur with the temporal adverbial tomorrow in (3-c) without resulting in ungrammaticality or at least a contradictory interpretation. therefore, the protagonist’s and the speaker’s perspective are conveyed in fid although the protagonist’s perspective is more prominent. 2.2. perspective in gesture. beside encoding perspective in spoken and written language, co-speech gestures have also been shown to encode perspective (mcneill 1992, among others). verbal utterances are often accompanied by gestures, which can either be manual, i.e., performed with the hands and potentially other body parts, or facial, i.e., performed with the face. these gestures are often synchronized with the verbal expressions they co-occur with. the stroke (= the core) of a gesture, for instance, is usually aligned with the nuclear accent of a word (loehr 2004, ebert et al. 2011). moreover, different alignment patterns of gesture and speech have been shown to have different semantic effects (ebert & ebert 2014). this claim has been experimentally validated by the study reported in ebert et al. (2022). previous research has distinguished different gesture types (for an overview, see mcneill 1992), among them iconic gestures. iconic gestures visually resemble a property of an object or action they illustrate. perspective is often encoded in iconic gestures. mcneill (1992) distinguishes between character viewpoint gestures (cvgs) and observer viewpoint gestures (ovgs, see also parrill 2010, stec 2012, among others). cvgs illustrate an event from a first-person perspective and the whole body is usually involved in the production of the gesture. ovgs, by contrast, illustrate an event as if observed from a distance (i.e., from a third-person perspective) and therefore only the hands are involved when producing the gesture. cvgs have been argued to be more informative than ovgs (beattie & shovelton 2002). examples are given in figures 2a and 2b, respectively. both are taken from the study reported in parrill (2010) where participants had to describe cartoon scenes as in figure 1 to friends who had not seen the clips. when describing the hopping movement of the skunk as shown in figure 1, speakers have several options to also incorporate gesture. the gesture shown in figure 2a is a clear instance of a cvg as the speaker depicts the skunk’s movement from a first-person and thus an internal perspective by imitating the skunk’s posture and also to a certain extent its facial expression. on the other side, the speaker shown in figure 2b produces a fairly clear instance of an ovg as they only trace the skunk’s trajectory and hopping with their index finger, thus adopting a more external, third-person perspective. there is a third type of viewpoint gesture, which occurs very infrequently, however. this gesture type encodes multiple viewpoints at the same time and has therefore been dubbed dual viewpoint gesture (parrill 2009). encoding multiple viewpoints in gesture seems to be less constrained than the occurrence of multiple viewpoints in speech since dual viewpoint gestures allow proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 413 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: cartoon scene of a skunk hopping across a room. taken from parrill (2010) (a) cvg used to depict the skunk in figure 1. (figure 3 in parrill 2010, p. 652) (b) ovg used to depict the skunk in figure 1. (figure 2 in parrill 2010, p. 651) figure 2: examples of a cvg and an ovg to depict the event shown in figure 1 for the presence of two (equally prominent) character viewpoints at the same time, which has not been attested for spoken or written language. interestingly, this is also possible in sign languages (e.g., maier & steinbach 2022). therefore, this might be a modality-specific feature. the copresence of a character’s and an observer’s viewpoint in a gesture has been argued to be possible although it often produces an ironic effect because the two viewpoints seem to compete with each other (mcneill 1992). moreover, dual cvgs are restricted to specific contexts at least for adults (mcneill 1992). more specifically, one of the cvgs always is a deictic gesture to the speaker’s body, representing the viewpoint of one character, and the body represents another character viewpoint. it can be noted, in sum, that the expression of multiple viewpoints is more liberal in gesture as opposed to speech. 2.3. interactions of perspective-taking in gesture and speech. expressing viewpoint in gesture is not entirely independent of the verbal material the gesture co-occurs with. parrill (2010), for example, has noted that an interdependence of linguistic, event, and discourse structure affects the choice of the gestural viewpoint. new information has been argued to co-occur with cvgs more frequently than with ovgs (mcneill 1992, parrill 2010). transitive utterances are more frequently accompanied by a cvg than by an ovg. moreover, events in which the speaker shows affect are also more likely to be accompanied by a cvg. it has furthermore been noted that events in which a trajectory is described are more likely to be accompanied by an ovg. crucially, however, viewpoint in gesture and speech have been argued to share a conceptual source (parrill proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 414 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 2009, 2010). in a similar vein, kita & özyürek (2003) present findings from crosslinguistic work suggesting that the execution and also the form of gestures depend on a language’s lexical and structural encoding. in other words, gesture execution and gesture form depend on the language of the speaker. if, for example, trajectory is encoded in a word (e.g., the english noun swing), a gesture co-occurring with this word is also more likely to encode trajectory. this constitutes a clear interaction of linguistic form and gesture execution, thus suggesting that there also is an interaction between linguistic and gestural viewpoint. the study reported in hinterwimmer et al. (2021) investigates the hypothesis that perspective in gesture and speech is by default aligned. the authors develop this hypothesis based on the assumption that gesture and speech convey a multimodal message which is planned by one central cognitive process. this message is then passed on to different communication channels (cf. mcneill 1992, de ruiter 1998, among others). moreover, as has been laid out above, viewpoint expressed in gesture and speech has the same conceptual source (parrill 2010) and perspective is an essential part of multimodal messages. finally, they claimed for the information conveyed in the two communication channels to be coherent by default (although gesture-speech mismatches can sometimes be useful, cf. goldin-meadow 1999), thus allowing them to straightforwardly derive the above-stated hypothesis. in order to test for their hypothesis, they constructed stimuli which either expressed the perspective of an individual participating in the event, i.e., a protagonist’s perspective, or the perspective of the speaker, i.e., a narrator’s perspective. these utterances were then paired with cvgs and ovgs. an example can be found in (4). underlined parts of an example indicate gesture-speech alignment. (4) a. narrator’s perspective: leon ist ein begeisterter sportler. als er sich neulich beim fußballspielen den ball erkämpfte, kickte er ihn sofort in richtung tor. + cvg/ovg ‘leon is an enthusiastic athlete. when he recently won the ball while playing soccer, he immediately kicked it in the direction of the goal.’ b. character’s perspective: leon spielte am wochenende fußball. nach einigem gerangel hatte er sich den ball erkämpft. toll, jetzt konnte er ihn direkt in richtung tor schießen! + cvg/ovg ‘leon played soccer on the weekend. after some scramble, he had finally won the ball. great, now he could directly kick it in the direction of the goal!’ cvg: speaker performs a kicking movement with their right leg and foot with an enthusiastic facial expression. ovg: speaker presses their index finger on their thumb, then releasing it quickly, thus imitating a kicking movement with their index finger. (hinterwimmer et al. 2021, p. 8) participants were presented either both versions of a stimulus in the narrator’s perspective, meaning that they saw videotaped versions of the same verbal utterance, once accompanied by a cvg and once accompanied by an ovg, or they saw both versions of the verbal stimulus in fid, i.e., in the protagonist’s perspective. hinterwimmer et al. (2021) used a forced-choice paradigm for the proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 415 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ study, meaning that participants had to choose the preferred version of the stimulus from the two available options. they predicted that stimuli expressing a protagonist’s perspective on the speech level should be preferred when they are accompanied by a cvg. when a stimulus expressed a narrator’s perspective, by contrast, the ovg was predicted to be preferred. contrary to their predictions, the results showed an overall preference for the cvg items regardless of the perspective expressed in the speech signal. there are two potential explanations for hinterwimmer et al.’s (2021) findings: i) there is no preference for perspective alignment in multimodal messages or ii) the default for linguistic and gestural perspective indeed is to be aligned, but intervening pragmatic factors can overwrite this preference. hinterwimmer et al. (2021) note that most of their stimuli contained transitive utterances as target sentences. as has been pointed out above, previous research has shown that cvgs tend to be preferred over ovgs in transitive utterances and that cvgs are more likely to be used when new information is expressed verbally (mcneill 1992, parrill 2010), which could thus have caused the overall preference for cvgs. moreover, cvgs are in general more informative than ovgs (beattie & shovelton 2002), which could be a further reason why cvgs were preferred over ovgs. the two types of viewpoint gestures differ in terms of size. while the whole body is usually involved in the production of a cvg, only hands and arms are involved in the production of an ovg (mcneill 1992). this size difference potentially makes cvgs more salient than ovgs (but see walter 2024 for tentative evidence against this claim), which could in turn also account for the observed cvg preference. in order to control for the potential pragmatic factors which might have blocked perspective alignment, the follow-up rating study reported in section 3 was conducted where only cvgs were used as gestures. using only cvgs controls for the aforementioned factors because they do not differ in terms of size and informativity, for example. this time, items were constructed where two protagonist’s perspectives were introduced on the speech level. one of them was prominent, the other one was not. in one condition, a cvg co-occurred with the speech signal that matched the prominent protagonist’s perspective on the speech level. in the other condition, a cvg was aligned with the verbal stimulus matching the non-prominent protagonist’s perspective. following hinterwimmer et al. (2021), it was hypothesized that this time, the condition with the cvg matching the prominent protagonist’s perspective on the speech level should be preferred as the factors potentially blocking perspective alignment discussed above were controlled for. 3. experimental study. 3.1. method. 3.1.1. participants. self-reported native speakers of german (n = 40) were recruited via prolific. they were naive with respect to the research question. 3.1.2. materials. for materials, 24 experimental items were constructed which were then videotaped. each item consisted of two sentences: the first one introduced an event and the second one further elaborated on that event. in the first sentence, a protagonist was introduced whose perspective was made prominent on the speech level and picked up again in the second sentence. a further, non-prominent protagonist’s perspective was introduced in the second sentence of each experimental item. in order to keep this perspective less prominent, this referent was always introduced by means of an indefinite dp (see meuser et al. to appear and meuser 2022 for experimental proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 416 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ evidence that referents introduced by indefinite dps are less prominent as perspective-takers than referents introduced by proper names). crucially, the prominent protagonist’s perspective was either introduced by means of a first-person pronoun or a proper name (factor referential expression). the utterances were accompanied either by a cvg matching the prominent protagonist’s perspective or by a cvg matching the non-prominent protagonist’s perspective (factor gesture). the study was thus of a 2x2 design. an example item is given in (5). (5) a. gestern abend ist mir etwas krasses passiert. ich war im park spazieren und auf einmal kam ein typ auf mich zu und hat mich ohne vorwarnung so heftig geschubstnot prominent, dass ich fast hingefallen wäreprominent, weil ich das gleichgewicht verloren habe. ‘yesterday evening something crazy happened to me. i was taking a walk in the park when suddenly some guy walked to me and nudgednot prominent me so strongly that i nearly fellprominent because i lost my balance.’ b. gestern abend ist paula etwas krasses passiert. sie war im park spazieren und auf einmal kam ein typ auf sie zu und hat sie ohne vorwarnung so heftig geschubstnot prominent, dass sie fast hingefallen wäreprominent, weil sie das gleichgewicht verloren hat. ‘yesterday evening something crazy happened to paula. she was taking a walk in the park when suddenly some guy walked to her and nudgednot prominent her so strongly that she nearly fellprominent because she lost her balance.’ prominent cvg: speaker is staggering backwards and flailing about. not prominent cvg: speaker performs a nudging gesture. in order to distract participants from the research question, the experimental items were interspersed with 25 unrelated fillers. in addition, there was a training session consisting of two items. 3.1.3. procedure. before the training session, participants were made familiar with the task by an introductory text. in this text, they were also instructed to pay attention to the audio as well as the videotape. moreover, they were informed about their data protection rights and gave informed consent. the questionnaire was created using sosci survey (leiner 2022), an online platform for creating questionnaires which can be used free of charge for academic purposes. the questionnaire was distributed via prolific using its pre-filtering functions to exclusively select for german native speakers as participants. the lists of the questionnaire were run independently in order to be able to filter out those participants who had completed a previous list of the questionnaire. this was done to prevent participants from participating multiple times. the items were split up according to a latin square design and evenly distributed onto four lists. the 25 fillers as well as the two training items were included on each list. experimental items and fillers occurred in a randomized order for each participant. in order to check the participants’ attention, they were asked to answer three questions about the videos after the final trial (e.g., they were asked what the hair color of the person they saw in the videos was). participants had to rate on a 7-point likert scale how natural they considered each utterance (1 = completely unnatural; 7 = completely natural). 3.2. predictions. based on the hypothesis that perspective in gesture and speech is aligned if there are no intervening pragmatic factors, the cvg representing the prominent protagonist’s perspective should always be preferred over the cvg representing the non-prominent protagonist’s perspective. moreover, since introducing a perspective by means of a first-person pronoun proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 417 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ makes this perspective more prominent compared to introducing it by means of a proper name (see bimpikou 2020 and saure et al. 2023 for experimental evidence for this assumption), the preference for the cvg matching the prominent protagonist’s perspective should be higher in the former case compared to the latter. thus, an interaction between the factors referential expression and gesture is predicted. 3.3. results. the data was analyzed using the r statistics software (r core team 2022). to test for significant effects, the results were analyzed using a cumulative link mixed effects model with the clmm() function in the r package ordinal (christensen 2023). for the analysis, an ordinal mixed effects model was chosen instead of a linear mixed effects model out of two reasons: first, linear mixed effects models require normally distributed data and second, they require the data to be measured at the interval level. both is questionable for likert scale data, making an ordinal mixed effects model the more adequate choice. the two factors were entered into the model as fixed effects using effect coding, that is the intercept represents the unweighted grand mean and the fixed effects compare the factor levels to each other. the full analysis script as well as the materials can be found here: https://osf.io/7hwjz/. the mean values and standard deviations (sds) are shown in figure 3. in general, there are only very subtle rating differences in the ratings for the cvg matching the prominent protagonist’s perspective (first-person pronoun: m = 5.43, sd = 1.53; proper name: m = 5.47, sd = 1.43) as well as for the cvg matching the non-prominent protagonist’s perspective (first-person pronoun: m = 5.39, sd = 1.47; proper name: m = 5.33, sd = 1.53). the results also show that there were no big rating differences in general between the cvg matching the prominent protagonist’s perspective and the cvg matching the non-prominent protagonist’s perspective. the ordinal mixed effects model corresponding to the data shown in figure 3 is given in table 1. unsurprisingly, neither main effects nor interactions are observable in the model output. figure 3: mean values and standard deviations for each condition. (abbreviations: 1p = firstperson pronoun, pn = proper name, prom = cvg from the prominent protagonist’s perspective, not prom = cvg from the not prominent protagonist’s perspective) proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 418 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ estimate std. error z value pr(> |z|) gesture .177 .121 1.472 .141 referential expression -.075 .121 -.619 .536 gesture:refrential expression .064 .241 .264 .792 table 1: ordinal mixed-effects model with mode and match as fixed effects and participants and items as random intercepts. formula: choiceo ∼ gesture * refexp + (1|case) + (1|item) significance codes: *** 0.001 | ** 0.01 | * 0.05 | . 0.1 3.4. discussion. in contrast to hinterwimmer et al.’s (2021) seminal study investigating alignment patterns of perspective in gesture in speech, the study reported in this paper controlled for intervening pragmatic factors which potentially block perspective alignment since in this study only cvgs were compared. however, the model output still does not confirm the hypothesis that perspective in gesture and speech is aligned in the absence of intervening pragmatic factors because this hypothesis straightforwardly translates to predicting an interaction between the factors referential expression and gesture. this interaction cannot be found in the data. furthermore, not even a main effect of gesture can be observed, indicating that the gesture matching the non-prominent protagonist’s perspective was equally preferred as the the gesture matching the prominent protagonist’s perspective. 4. general discussion and conclusion. the study reported in hinterwimmer et al. (2021) (cf. section 2.3 of this paper) for the first time investigated alignment patterns of gestures and speech. based on i) findings that a multimodal message is planned by one underlying cognitive process (e.g., de ruiter 1998), ii) findings that viewpoint in gesture and speech have the same conceptual source (parrill 2010), and iii) the assumption that information conveyed in the two communication channels is by default coherent, the authors hypothesized that perspective in gesture and speech is by default aligned. they investigated this in a forced-choice study comparing cvgs and ovgs. the results, however, showed an overall cvg preference although this preference was slightly (but not significantly) smaller in the condition where the narrator’s perspective was prominent on the speech level. based on this slightly smaller preference, they argued that there might have been pragmatic factors that blocked the hypothesized perspective alignment. examples for these pragmatic factors are the transitivity of the target sentences in their study (transitive events more likely evoke cvgs, cf. parrill 2010) or salience differences between the two gesture types (but see walter 2024 for results suggesting otherwise). they therefore concluded with the weakened hypothesis that perspective in gesture and speech is aligned, but only if there are no intervening pragmatic factors blocking the alignment. the study reported in this paper investigated hinterwimmer et al.’s (2021) weakened hypothesis by means of a rating study controlling for the aforementioned pragmatic factors which might have blocked the perspective alignment in their study. in this study, two protagonist’s perspectives were introduced on the speech level, one of them being prominent and the other one being not prominent. these verbal stimuli were then either aligned with a cvg matching the prominent protagonist’s perspective or a cvg matching the non-prominent protagonist’s perspective. it was hypothesized that the cvg matching the prominent protagonist’s perspective should be preferred. proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 419 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ this was not borne out by the results, however. taking together the results reported in hinterwimmer et al. (2021), both studies point toward rejecting the hypothesis as both times no alignment preference could be obtained from the data. instead, it seems that multiple perspectives can be expressed relatively freely in gesture and speech. it remains an open question, however, if there are any constraints on which perspectives can be simultaneously expressed in the two modalities. still, it might be a little premature to reject the perspective alignment hypothesis as will be laid out below in further detail. the results of the studies are somewhat surprising, since there is empirical evidence in favor of a preferred perspective alignment from sign languages. in a study with signers of german and turkish sign language, for example, it has been found that signers tend to adopt a character’s perspective when producing a handling classifier, i.e., a classifier showing how an object was handled, and an observer’s perspective for entity classifiers, that is classifiers denoting an entity from an outside perspective (özyürek & perniss 2011). moreover, an overall preference for adopting a character’s perspective for signers of german and turkish sign language was observed in this study. this could in principle at least explain the cvg preference observed by hinterwimmer et al. (2021). the preference to adopt a character perspective could then be a specific property of the visual modality also in spoken languages. a potential explanation for this character viewpoint preference is as follows: it has been noted at least for signers that they prefer depicting over describing when reporting an event (engberg-pedersen 1993, özyürek & perniss 2011). since adopting a character viewpoint in gesture allows for richer depictions compared to adopting an observer viewpoint as the whole body is involved when adopting the former viewpoint (cf. mcneill 1992), the character viewpoint preference might come about because of a general preference for depicting over describing when reporting events. in general, the visual modality is more suitable than the acoustic modality to depict events. assuming that the preference to depict when reporting events also holds for spoken languages, the observed preference for cvgs is easily explained. this, in turn, can also be transferred to the findings of the study reported in this paper: although there potentially is a preference for perspective alignment also in spoken languages, both cvgs used in the study were equally depictive. it is thus possible that the experimental items were not suitable to test for our hypothesis. in addition, the preference to adopt a character perspective in the visual modality and to therefore prefer cvgs in spoken language might be so strong that perspective alignment is only experimentally testable in cases where the perspective of a protagonist is made extremely prominent on the speech level. one of these cases are so-called be like-constructions where not only words, but also actions and other non-linguistic behavior can be quoted under a demonstrational account to quotation (clark & gerrig 1990, davidson 2015). (6) john was like hurrying to get to the bus station on time. + cvg depicting running be like-constructions thus have a highly depictive component while at the same time making a protagonist’s perspective very prominent on the level of speech as they directly quote that protagonist’s speech and/or actions. intuitively, an alternation of (6) with an ovg depicting john’s running instead of a cvg seems less natural than (6). this indicates that here, perspective alignment holds. before ultimately rejecting the hypothesis of a preference for perspective alignment, this should be tested experimentally. we leave this to future research. proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 420 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ references anand, pranav & andrew nevins. 2004. shifty operators in changing contexts. in robert b. young (ed.), proceedings of semantics and linguistic theory (salt) 14, 20–37. ithaca, ny: clc publications. 10.3765/salt.v14i0.2913. anderson, carolyn j. 2019. tomorrow isn’t always a day away. in m. teresa espinal, elena castroviejo, manuel leonetti, louise mcnally & cristina real-puigdollers (eds.), proceedings of sinn und bedeutung 23, 37–56. barcelona, spain: universitat autònoma de barcelona. beattie, geoffrey & heather shovelton. 2002. an experimental investigation of some properties of individual iconic gestures that mediate their communicative power. british journal of psychology 93(2). 179–192. 10.1075/gest.1.2.03bea. bimpikou, sofia. 2020. who perceives? who thinks? anchoring free reports of perception and thought in narratives. open library of humanities 6(2). https://doi.org/10.16995/olh.484. christensen, rune h. b. 2023. ordinal—regression models for ordinal data. https://cran. r-project.org/package=ordinal. r package version 2023.12-4. clark, herbert h. & richard j. gerrig. 1990. quotations as demonstrations. language 66(4). 764–805. 10.2307/414729. davidson, kathryn. 2015. quotation, demonstration, and iconicity. linguistics and philosophy 38(6). 477–520. 10.1007/s10988-015-9180-1. ebert, cornelia & christian ebert. 2014. gestures, demonstratives, and the attributive/referential distinction. talk given at semantics and philosophy in europe 7. ebert, cornelia, stefan evert & katharina wilmes. 2011. focus marking via gestures. in ingo reich, eva horch & dennis pauly (eds.), proceedings of sinn und bedeutung 15, 193–208. saarbrücken, germany: university of saarland. ebert, cornelia, giovanna pirillo & sebastian walter. 2022. the role of gesture-speech alignment for gesture interpretation. in sam featherston, robin hörnig, andreas konietzko & sophie von wietersheim (eds.), proceedings of linguistic evidence 2020: linguistic theory enriched by experimental data, 65–77. tübingen, germany: university of tübingen. engberg-pedersen, elisabeth. 1993. space in danish sign language: the semantics and morphosyntax of the of space in a visual language. hamburg, germany: signum. goldin-meadow, susan. 1999. the role of gesture in communication and thinking. trends in cognitive sciences 3(11). 419–429. 10.1016/s1364-6613(99)01397-2. harris, jesse a. 2012. processing perspectives. amherst, ma: university of massachusetts amherst dissertation. harris, jesse a. & christopher potts. 2009. perspective-shifting with appositives and expressives. linguistics and philosophy 32(6). 523–552. 10.1007/s10988-010-9070-5. hinterwimmer, stefan. 2017. two kinds of perspective taking in narrative texts. in dan burgdorf, jacob collard, sireemas maspong & brynhildur stefánsdóttir (eds.), proccedings of semantics and linguistic theory (salt) 27, 282–301. university park, md: university of maryland. 10.3765/salt.v27i0.4153. hinterwimmer, stefan. 2024. accounts of perspective taking in narrative. language and linguistics compass 18(3). e12517. 10.1111/lnc3.12517. hinterwimmer, stefan, umesh patil & cornelia ebert. 2021. on the interaction of proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 421 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ gestural and linguistic perspective taking. frontiers in communication 6. 1–15. 10.3389/fcomm.2021.625757. kita, sotaro & asli özyürek. 2003. what does cross-linguistic variation in semantic coordination of speech and gesture reveal? evidence for an interface representation of spatial thinking and speaking. journal of memory & language 48(1). 16–32. 10.1016/s0749-596x(02)00505-3. leiner, daniel j. 2022. sosci survey (version 3.3.14a). available at https://www.soscisurvey.de. loehr, daniel p. 2004. gesture and intonation. washington, dc: georgetown university dissertation. maier, emar & markus steinbach. 2022. perspective shift across modalities. annual review of linguistics 8(1). 59–76. 10.1146/annurev-linguistics-031120-021042. mcneill, david. 1992. hand and mind: what gestures reveal about thought. chicago, il: university of chicago press. meuser, sara. 2022. how free is free indirect discourse? empirical approaches to the anchoring mechanisms of perspective-taking. cologne, germany: university of cologne dissertation. meuser, sara, maximilian hörl & stefan hinterwimmer. to appear. perspective-taking and protagonist prominence: an empirical approach to the role of local and global prominence. to appear in discourse and dialogue. özyürek, asli & pamela m. perniss. 2011. event representations in signed languages. in event representations in language and cognition, 84–107. cambridge university press. parrill, fey. 2009. dual viewpoint gestures. gesture 9(3). 271–289. 10.1075/gest.9.3.01par. parrill, fey. 2010. viewpoint in speech-gesture integration: linguistic structure, discourse structure, and event structure. language and cognitive processes 25(5). 650–668. 10.1080/01690960903424248. plank, frans. 1986. über den personenwechsel und den anderer deiktischer kategorien in indirekter rede. zeitschrift für germanistische linguistik 14(3). 284–308. 10.1515/zfgl.1986.14.3.284. r core team. 2022. r: a language and environment for statistical computing. r foundation for statistical computing vienna, austria. https://www.r-project.org/. de ruiter, jan p. 1998. gesture and speech production. nijmegen, the netherlands: university of nijmegen dissertation. saure, christopher, stefan hinterwimmer & anna pia jordan-bertinelli. 2023. an experimental investigation of the interaction of narrators’ and protagonists’ perspectival prominence in narrative texts. zeitschrift für sprachwissenschaft 42(2). 341–372. schlenker, philippe. 2004. context of thought and context of utterance: a note on free indirect discourse and the historical present. mind & language 19(3). 279–304. 10.1111/j.14680017.2004.00259.x. stec, kashmiri. 2012. meaningful shifts: a review of viewpoint markers in co-speech gesture and sign language. gesture 12(3). 327–360. 10.1075/gest.12.3.03ste. walter, sebastian. 2024. the at-issue status of viewpoint gestures: evidence for gradient atissueness. in geraldine baumann, daniel gutzmann, jonas koopman, kristina liefke, agata renans & tatjana scheffler (eds.), proceedings of sinn und bedeutung 28, 943–960. bochum, germany: ruhr university bochum. 10.18148/sub/2024.v28.1171. proceedings of elm 3: 411-422, 2025 sebastian walter and stefan hinterwimmer: an experimental investigation of perspective alignment in gesture and speech. 422 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ proportions vs. cardinalities: comparative ambiguities and the covid pandemic elsi kaiser * abstract. this paper reports two psycholinguistic experiments on quantity comparatives and superlatives that are potentially ambiguous between cardinal and proportional readings. by using statements about covid cases and vaccination numbers as a naturalistic context with real-world relevance, this work furthers our understanding of what happens in linguistic environments where multiple measure functions are available – what modulates the choice between them? the results provide new evidence that comparatives and superlatives can refer to scales ranging over degrees of proportion, in addition to degrees of cardinality. furthermore, this experimental evidence points to a preference for cardinal interpretations, but also shows that this preference is not rigid and can be weakened in favor of proportional readings by semantic factors – including considerations potentially related to stage vs. individual-level differences – and by certain linguistic forms. keywords. comparatives; superlatives; experimental semantics; proportional quantifiers; degree semantics; covid-19; stage-level and individual-level predicates 1. introduction: cardinal and proportional scales. in degree-based semantics, comparatives express relations between degrees on a scale (e.g. creswell 1977, von stechow 1984, beck 2011). for example, example (1) is likely to be true if we compare cardinalities of people. philadelphia is a much bigger city than bryn mawr (a small town in the philadelphia suburbs). thus, the statement in (1) is likely to be true if we are counting the raw numbers of people who know their neighbors in philadelphia vs. bryn mawr. on this reading, we are assuming a scale ranging over degrees of cardinality. but, interestingly, the ‘reverse’ in (2) is also likely to be true – if we compare proportions. since bryn mawr is a smaller town, and assuming that people in smaller places are more likely to be friendly with their neighbors, it’s likely that a larger proportion of people in bryn mawr than in philadelphia know their neighbors. now, we are dealing with a scale that ranges over degrees of proportions. examples (1-2) are adapted from partee (1989). (1) cardinal reading: more residents of philadelphia than bryn mawr know their neighbors. |philadelphiaknow_neighbor| > |bryn mawrknow_neighbor| (2) proportional reading: more residents of bryn mawr than philadelphia know their neighbors. |bryn mawrknow_neighbor| / |bryn mawrpopulation| > |phillyknow_neighbor| / |phillypopulation| these kinds of cardinal-proportional ambiguities also occur with quantifiers such as few and many. partee (1989) noted that sentences such as ‘few aspens burned’ and ‘many aspens * many thanks are due to the elm 2 audience for helpful comments and feedback and to roumyana pancheva and alexis wellwood for comments on earlier versions of this work. i would also like to thank haley hsu and madeline rouse, as well as claire post and deborah ho, for their help with stimulus creation and experiment set-up. i gratefully acknowledge funding from the social science research council and the usc zumberge special solicitation: epidemic& virus-related research and development award. author: elsi kaiser, university of southern california (emkaiser@usc.edu) proceedings of elm 2: 142-153, 2023 c©2023 elsi kaiser published by the lsa with permission of the author(s) under a cc by license. 142 https://doi.org/10.3765/elm https://www.elm-conference.net/ burned’ are ambiguous between cardinal readings (i.e., a large/small number of aspens burned) and proportional readings (i.e., a large/small proportion of the (contextually relevant) aspens burned). see also westerståhl (1985), herburger (1997), romero (2015), solt (2018) and bale & schwarz (2020), and many others for related discussion. though both cardinal and proportional readings have been argued to be available, questions remain about how to formally capture the semantics of these readings. some recent theoretical work has explored the idea of using underspecified measure functions (see e.g. bale & schwarz 2020, bale & barner 2009, see also wellwood 2015 and subsequent work). for example, bale & schwarz discuss examples involving what they call ‘contextual proportionality,’ where the measure function is not fully specified by the conventional meaning of the sentence and instead is contextually determined. it’s also worth noting that, as bale & schwarz point out, the availability of cardinal interpretations can make it hard to detect the existence of proportional interpretations. indeed, intuitively, cardinal readings are preferred (‘easier to get’). this raises questions both about the strength/robustness of this cardinality bias (can it be weakened or even overcome?), and about the factors that might modulate the preference for one reading over the other. i take a closer look at these issues next. 1.1. cardinality bias. broadly speaking, prior work seems to assume that cardinal readings are preferred over proportional readings (e.g. solt 2018, bale & schwarz 2020), but – to the best of my knowledge – this has not been systematically experimentally tested. thus, this paper reports two studies that aim to assess, by means of two psycholinguistic studies, whether one reading is preferred over the other, in contexts where both cardinal and proportional are contextually available. from a processing point-of-view, it seems that a dispreference for proportional readings would not be unexpected, due to their greater complexity: proportional readings depend on both the numerator and the denominator, whereas cardinal readers essentially only make reference to the numerator. this leads us to expect a cardinality bias, a preference to interpret ambiguous comparatives as comparing cardinalities, not proportions (other things being equal). 1.2. individual vs. stage-level predicates. in addition to the cardinality bias, prior work suggests that cardinal vs. proportional readings (at least with certain quantifiers) are constrained by semantic considerations about predicate type (e.g. partee 1989, huettner 1984, solt 2018, see also milsark 1977), in particular the distinction between stage-level predicates (describing transient properties, e.g. is feverish) vs. individual-level predicates (describing more permanent properties, e.g. has a college degree, carlson 1977). more specifically, it has been noted that, when combined with individual-level predicates, many and few seem to only allow (or perhaps very strongly prefer) the proportional reading, not the cardinal reading. although prior work does not specifically discuss comparatives or superlatives of the specific types tested in the experiments in this paper, if these observations are more broadly relevant, they lead us to expect that individual-level predicates might favor proportional interpretation more than stage-level predicates. in other words, the hypothesized cardinality bias might be weakened when we are comparing individual-level properties. the stage-level vs. individual-level distinction corresponds well to two kinds of statements that were highly prevalent during the height of the covid pandemic, namely statements about covid cases/infections and statements about vaccination numbers. while having covid can be regarded as a non-permanent, transient, stage-level property, being vaccinated is more indiproceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 143 https://doi.org/10.3765/elm https://www.elm-conference.net/ vidual-level, permanent property.1 indeed, the prediction that individual-level predicates favor proportional interpretations more than stage-level predicates seems to match how information about covid cases and vaccination numbers was communicated during the height of the pandemic (at least in the u.s.): impressionistically, it seems that covid cases were typically reported in the news and in public health communication using raw numbers (cardinal) or per 100,000 (proportional), whereas vaccination numbers appeared to be mostly reported using percentages (proportional).2 thus, the second aim of the present work is to explore the link between stagevs. individual-level predicates and the likelihood of cardinal vs. proportional readings through the lens of pandemic-related comparatives. this difference in how covid case numbers and vaccination numbers are conceptualized is also revealed by how these two types of information are graphically reported in (u.s.) public health communications: while covid cases are often graphed in a non-cumulative way (showing temporary spikes in case numbers during surges/waves), vaccination numbers tend to be graphed in a cumulative manner (as gradually increasing curves). this indicates, again, that we are more likely to regard having covid as a more temporary state (stage-level) and being vaccinated as a more permanent property (individual-level). having said this, it is important to acknowledge that other factors beyond the stagevs. individual-level distinction may also turn out to be at play, as discussed below. thus, the studies here represent an initial step that should be complemented by further work. 1.3. predictions. the two experiments reported here explore two (non-mutually exclusive) predictions: first, the cardinality bias, which predicts that cardinality readings are generally preferred over proportional readings in quantity comparatives (experiment 1) and in superlatives (experiment 2). second, the experiments also test the semantic factors hypothesis, according to which the availability of cardinality and proportional interpretations can be modulated by semantic factors (in ways that may be related to the individual-/stage-level distinction). to test these predictions, i used sentences about (i) covid cases/infections (stage-level), and about (ii) vaccinations/vaccinated people (more individual-level). 1.4. the covid pandemic as a natural experiment. to investigate whether ambiguous comparatives and superlatives about covid cases and vaccinated people exhibit a preference for cardinal readings over proportional readings, and to see whether this is modulated by semantic factors (transient vs. more stable properties), two experiments were conducted. the covid pandemic provides a natural context for investigating the cardinal-proportional ambiguity, because it is a distinction that becomes very relevant when talking about covid infections or vaccinations. this is illustrated by (3), a (simplified version of a) naturally-occurring example from the internet, where person a is presumably assuming a cardinal interpretation and doubts its veracity, and b seems to be trying to explain that the relevant reading is the proportional one. thus, the pandemic provides a meaningful, naturalistic context for experiments on this topic. (3) confusion between cardinal vs. proportional (simplified from www) person a: alaska has more covid than california…riiiight. 1 i think it is fair to say that this contrast stands despite the existence of long covid and the fact that vaccine effectiveness wanes over time. this is because an acute covid infection is (as far as we know) temporally constrained while the state of having received a covid vaccination is permanent. 2 these are impressionistic assessments. a quantitative corpus study has not been conducted, as far as i know. proceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 144 https://doi.org/10.3765/elm https://www.elm-conference.net/ person b: no, the percentage goes by their population individually (…) yeah california is bigger buuut the percentages are only going off each states numbers (…) they aren’t counting people, only percentages of those people 2. experiment 1: comparatives. 2.1. participants. data from 139 self-reported adult native u.s. english speakers, recruited via amazon mturk, is reported. participation took place remotely over the internet. 2.2. design, materials and predictions. to test the interpretation of comparatives, eight pairs of covid information dashboards were created, mimicking the dashboards used in the u.s. to report covid cases and vaccination rates during the height of the pandemic. each dashboard provided information for an imaginary u.s. county. each dashboard reported the raw number of new covid cases (cardinal information), the per 100,000 covid case rate (proportional information), the raw number of fully vaccinated people (cardinal information), and the percent of fully vaccinated people (proportional information). thus, both proportional and cardinal information are saliently available on the dashboards. this means that any differences in people’s willingness to consider cardinal vs. proportional readings cannot be straightforwardly attributed to the proportional reading being very ‘low salience’ or hard to access, given that it is explicitly depicted on the dashboards. the dashboards also mentioned the number of covid tests done and the positivity rate, but this information was blurred out as it was not relevant for this study. (the study was designed and implemented before younger children were eligible for vaccination in the u.s. and thus the information about vaccination rates excludes children under age 12.) figure 1: screenshot of an example display from experiment 1 the dashboards, which were presented to participants in pairs as shown in figure 1, were designed so that one county had higher raw numbers of covid cases (or higher raw numbers of vaccinated people), and the other one had a higher proportional covid rate (or higher percentproceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 145 https://doi.org/10.3765/elm https://www.elm-conference.net/ age of vaccinated people). thus, the cardinal vs. proportional readings are truth-conditionally distinct. (the magnitude of the differences in covid cases and vaccination numbers between the two counties varies across items). on each screen, below the pair of dashboards, participants saw a critical sentence with two blanks where the names of the counties should be. their task was to type the missing information into the textbox below the sentence. a second textbox (not shown above) was also included, asking participants to explain what information they used to fill in the blanks. this was done to keep participants on task and to provide us with an additional means of checking participants’ responses if their answers were unclear. due to the truth-conditionally distinct design, participants’ responses reveal whether they opted for a cardinal or a proportional interpretation. for example, for the display and the sentence in figure 1 (there are more covid cases in ___ than ___ ), putting teakley county in the first blank and hawkton county in the second blank indicates a cardinal interpretation, whereas putting hawkton county in the first blank and teakley county in the second blank indicates a proportional interpretation. the study tested eight different linguistic frames, four involving covid cases and four involving vaccination. four of these function as control conditions, because the lexical semantics of the sentence (e.g. use of phrases such as number of covid cases vs. rate of covid cases) signal what degrees the scale ranges over, i.e., provides information about the intended measure function. these are shown in (4) for the cardinal controls and (5) for the proportional controls. (4) cardinal control the number of covid cases is higher in ___ than in ___ the number of fully vaccinated people is higher in ___ than in ___ (5) proportional control the rate of covid cases is higher in ___ than ___ the vaccination rate is higher in ___ than in ___ i also tested the more ambiguous structures in (6), where there are no lexical items indicating what degrees the scale ranges over. i refer to these as the more + noun conditions. it’s worth noting that both cardinal and proportional readings are contextually available; both kinds of numbers are on the dashboards. thus, we can ask, in the absence of lexical semantic cues, which interpretation do people prefer, and does this differ for covid cases vs. vaccinated people? (6) more + noun condition there are more covid cases in ___ than ___ there are more fully vaccinated people in ___ than ___ finally, two exploratory conditions were also included. example (7) shows the truncated noun and the truncated adjective conditions. in these conditions, the comparative only mentions more followed by the bare noun covid or the adjective vaccinated, without mentioning ‘cases’ or ‘people.’ thus, the link between linguistic form and the measure function is even less clear. (7) truncated noun and adjective conditions: ___ has more covid than ___ [truncated with noun] ___ is more vaccinated than ___ [truncated with adjective] before continuing, it’s important to address a potential concern regarding the acceptability of these truncated structures. while they may sound marked to some speakers, our corpus invesproceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 146 https://doi.org/10.3765/elm https://www.elm-conference.net/ tigations indicate that they occur in natural use, in both formal/official and more informal contexts. some examples from the internet are provided in (8-9). (8) a. canada is more vaccinated than the uk. b. the teaching staff at woodstock school district 200 also is more vaccinated than its other workers (9) a. lubbock now has more covid than colorado’s largest city b. state of ny has more covid than my town. d. but either way, america has more covid than any of the countries he put travel bans on, so how well did they work? d. “remember people were making fun of america, that we didn’t do it right, we had more covid than everybody else,” he said. in the truncated conditions (7), the thing being compared (cases or people) is not explicitly mentioned, and furthermore, due to the structure of the sentence, the placename acts as a proxy to refer to its inhabitants (for further discussion of meaning transfer, see e.g. nunberg 1995). a place does not get covid or get vaccinated, its inhabitants do. despite this similarity, the truncated conditions with nouns and adjectives have different properties: the truncated conditions with nouns (__ has more covid than __) could possibly be construed as covert plurals with an elided plural noun (__ has more covid cases than __), or perhaps with covid construed as a mass noun, akin to ‘x has more water than y’. in both cases, there is no explicit reference to individual cases, and thus we might expect the cardinal reading to be relatively less available than in the more + noun conditions in (6). when we consider the truncated conditions with adjectives (__ is more vaccinated than __), it’s worth noting that the comparatives with placenames-standing-in-for-people differ from ‘regular’ adjectival comparatives where we are comparing individual people (10-11). when we compare two individual people to each other using a truncated adjective comparative (10a), we get a reading where the two people differ in terms or their degree of vaccination (e.g. kate has only had one shot of a two-shot vaccine and lisa has had both shots, or kate has not yet received a booster shot but lisa has). unsurprisingly, use of ‘fully’ is infelicitous in this case (10b). (10) adjectival comparatives with names a. lisa is more vaccinated than kate. b. lisa is more (*fully) vaccinated than kate (11) adjectival comparatives with placenames a. california is more vaccinated than alaska. b. california is more (okfully) vaccinated than alaska. however, when we are dealing with two placenames in a truncated adjective comparative (11), and the placenames stand in for their inhabitants (i.e., a plurality), we now instead get cardinal and proportional readings with sentences like (11a) and use of ‘fully’ is felicitous (since we are no longer comparing degrees of vaccination) (11b). this contrast between (10) and (11) suggests that, semantically, in the truncated adjective comparatives we are comparing two pluralities, which is indeed what we expect if the placename acts as a proxy for referring to the inhabitants of that place. however, this is not explicitly reflected in the sentence (the placenames are morphologically singular) and in fact there is no noun mentioned at all in the truncated adjective comparatives. thus, intuitively, this condition is perhaps the one where accessing a cardinality scale ranging over numbers (of vaccinated people) might be the ‘hardest.’ if this line proceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 147 https://doi.org/10.3765/elm https://www.elm-conference.net/ of thinking is on the right track, we expect the cardinal construal to be less likely in the truncated adjective condition than in the more + noun comparatives about vaccination (6), because the latter explicitly mention ‘people.’ in sum, properties of the (admittedly exploratory) truncated conditions in general lead us to expect a weakening of the cardinality bias, which might be clearest in the truncated adjective condition. more broadly, comparing the four different linguistic frames can provide insights into how linguistic packaging interacts with the predicted cardinality bias and predicate type effects. 2.3. differences between the covid conditions and vaccination conditions. it’s important to acknowledge that the sentences comparing covid cases and vaccinated people differ not only in terms of their stagevs. individual-level properties, but also in other ways. some of these are shown in table 1. covid cases vaccinated people property stage level individual level nominal cases allows double-counting people no double-counting denominator out of 100,000 out of population (percentage) table 1: semantic properties of comparatives involving covid cases vs. vaccinated people the nominal itself is different in the two types, cases with covid comparatives and people with vaccine comparatives. i suggest that the noun cases, but not the noun people, is what krifka (1990) calls a phase noun. phase nouns differ from non-phase nouns in allowing doublecounting. krifka shows by this by contrasting the nouns passengers and persons. consider a sentence like ‘two million passengers/persons passed through this airport in the last year.’ with persons, only an object-related reading is available (i.e., two million different individuals), whereas passengers also allows for an event-related reading (i.e., there were two million events of individuals passing through the airport, but this might only involve one million different people if everyone flew twice). i suggest that cases vs. people shows the same asymmetry. i do not provide a detailed discussion here, but this distinction might facilitate a cardinal measure function with the phase noun cases, as compared to the non-phase noun people. another distinction concerns the denominator relevant for the proportional reading. covid cases, when conceptualized proportionally, are typically reported out of 100,000 (at least in the u.s.), whereas vaccination numbers, when conceptualized proportionally, are typically reported as a percentage of the (relevant) population. these same denominators were used in the studies reported here (as can be seen in the dashboards in figure 1), in order to make the stimuli maximally naturalistic. however, it could be argued that conceptualizing proportional information in terms of percentages is more natural and ‘easier,’ and thus the proportional reading may be more salient than conceptualizing proportional information in terms of the arguably less widely-used ‘per 100,000’ metric. thus, the proportional reading may seem more artificial and thus less accessible with covid cases per 100,000 than with the percentage of fully vaccinated people. ultimately, there are multiple reasons, not just the stage vs. individual level distinction, why comparatives involving covid cases may be less likely to elicit proportional readings (and conversely, more likely to elicit cardinal readings) than comparatives involving vaccinated people. proceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 148 https://doi.org/10.3765/elm https://www.elm-conference.net/ admittedly, comparing two things differing along multiple dimensions is complicated and raises the question of why would one opt for this pairing. i opted to test comparatives about covid cases vs. vaccinated people for two reasons. first, i wanted to test how people interpret comparatives about information that was highly relevant during the height of the pandemic – which led us to look at infections and vaccination – and that was expressed in a natural/typical way. second, as a first step, i wanted to see whether we can detect any semantic effects at all – even if we do not know what precise factor is responsible, can we detect asymmetries in the likelihood of cardinal vs. proportional construals for covid cases vs. vaccinated people? in other words, my aim is to see whether we can detect any semantic effects on people’s interpretations in contexts where multiple measure functions are in principle available. this information can provide a foundation for future work seeking to systematically disentangle the different ways in which these two domains differ, to see which differences are ultimately responsible for the likelihood of opting for cardinal vs. proportional readings. thus, this work is best regarded as a first step testing whether this conglomeration of semantic factors matters. if we find evidence that comparatives about covid cases are interpreted differently than comparatives about vaccination numbers, this discovery would pave the way for a subsequent, more systematic comparison (perhaps in a non-covid domain) of how factors such as stage vs. individual level, phase nouns vs. non-phase nouns and different denominators contribute. 2.4. procedure. in experiment 1, participants saw displays like figure 1 and typed into the textbox the names of the counties in the order that they should go in the blanks. the study was untimed, i.e. participants were not under time pressure. the dashboards and the critical sentence were shown on the same screen so there was no memory load. 2.5. results. the results for experiment 1 are shown in figure 2. as expected, the control conditions mentioning rate elicit a predominance of proportional interpretations (mostly light grey in the top two bars) and the control conditions mentioning number elicit a predominance of cardinal interpretations (mostly dark grey in the next two bars). furthermore, we also find a difference between covid cases vs. vaccinated people: there are more cardinal responses with comparatives about ‘covid cases’ than with ‘vaccinated people’ (p’s < 0.05), even in these largely unambiguous control conditions. this provides support for the semantic factors hypothesis. turning to the more + noun conditions (fifth and sixth bars from the top), participants mostly provided cardinal responses (mostly dark grey), corroborating the predicted cardinality bias. furthermore, we again find a semantic effect, with more cardinal responses with ‘covid cases’ than ‘vaccinated people’ (p’s < 0.05). finally, looking at the truncated conditions (two bottom bars), it’s clear that while the truncated noun condition with ‘covid’ still exhibits a cardinality bias, such a bias is no longer present in the truncated adjective condition with ‘vaccinated’ which elicits more proportional than cardinal responses. this fits with our hunch that this might be the condition where the accessing a cardinality scale might be the hardest because there is no explicit mention of a plurality. furthermore, the truncated noun ‘covid’ condition (eighth bar) elicits fewer cardinal responses than the more + covid conditions (sixth bar), even though both are about covid. in other words, the cardinal bias we saw in the more + covid conditions is weakened in the truncated versions where no explicit reference is made to cases. proceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 149 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2: exp1 results (cardinal vs. proportional readings with comparatives; error bars +/1 se) 2.6. discussion. overall, these results provide clear support for the hypothesized cardinality bias and the semantic factors hypothesis which posits that the strength of the cardinality bias can be modulated by factors related to the target of comparison, with comparisons involving covid cases eliciting more cardinal interpretations than comparisons involving vaccinated people. furthermore, the results show that linguistic form also plays a role. in the more + noun conditions, which lack explicit lexical semantic information signaling the measure function, participants are less likely to opt for proportional readings than in the proportion control conditions that explicitly signal, by means of the word rate, that we are dealing with a measure function ranging over proportions. in addition, when no explicit reference is made to cases or to people, in the exploratory truncated noun and adjective conditions, the cardinality bias is weakened even more. this shows that linguistic packaging plays a key role in modulating the ease of accessing, or perhaps the salience of, proportional measure functions relative to default cardinal measure functions. 3. experiment 2: superlatives. the second experiment tests whether the effects observed in experiment 1 extend to superlatives, arguably a cognitively more complex situation because three different counties are being compared. based on semantic theory, we expect the results from experiment 1 to replicate. 3.1. participants. data for 129 adult native u.s. english speakers, recruited via amazon mturk, is reported for experiment 2. none had participated in experiment 1. again, participation took place remotely over the internet. 3.2. design and materials. the design and materials were the same as for experiment 1, except that now participants saw three dashboards on each screen and filled in single blanks in superlative sentences of the forms shown in (12-15). the different sentence types (cardinal control, proportional control, most + noun, and truncated noun and truncated adjective) were kept as similar to experiment 1 as possible, but adapted into superlatives. (12) cardinal control ___ has the highest number of covid cases ___ has the highest number of fully vaccinated people proceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 150 https://doi.org/10.3765/elm https://www.elm-conference.net/ (13) proportional control ___ has the highest rate of covid cases ___ has the highest vaccination rate (14) more + noun condition ___ has the most covid cases ___ has the most fully vaccinated people (15) truncated noun and adjective conditions ___ has the most covid [most + noun] ___ is the most vaccinated [most + adjective] as in experiment 1, some readers may have questions about the naturalness of the truncated forms. based on our corpus investigations, they occur in naturalistic texts. examples from the internet are in (16-17). (16) a. the figure below illustrates that among midwestern counties the well “wired” urban centers have the most covid-19, with “unwired” rural hinterlands and smaller com munities relatively unscathed. b. the richest and one of the most modern countries in the world has the most covid and it is increasing exponentially c. also, isn't it coincidental that the most tested populations have the most covid, and those unvaccinated untesting countries have almost no cases? (17) a. out of all the zip codes in douglas county, 68007, is the most vaccinated. b. she said the older population is the most vaccinated. figure 3: exp2 results (cardinal vs. proportional readings with superlatives, error bars +/1 se) 3.3. procedure. the procedure was the same for experiment 1, except that participants now saw three dashboards per screen and saw superlative sentences with only one blank (12-15). 3.4. predictions. the predictions are the same as for experiment 1, assuming that the effects extend to a more cognitively complex context that requires comparing three elements. proceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 151 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.5. results and discussion. the experiment 2 results are in figure 3, which shows that the key results from experiment 1 replicate with superlatives (though note that the rate of cardinal responses in the proportional controls mentioning covid cases, in the second bar from the top, is essentially at chance in experiment 2; importantly, this is still much lower than the rate of cardinal responses in the cardinal controls and thus largely fits with what we’ve seen so far.) 4. general discussion. this paper reports two psycholinguistic studies on quantity comparatives and superlatives that are potentially ambiguous between cardinal and proportional readings, with the aim of improving our understanding of what happens in linguistic environments where multiple measure functions are available. the studies test how likely participants are to opt for a cardinal or a proportional measure function, and what factors modulate these preferences. the covid-19 pandemic provides a naturalistic context for exploring these kinds of questions, because the cardinal vs. proportional distinction is very important when considering information about covid cases and vaccination rates. for example, if someone says that small town a with less than ten thousand inhabitants has more covid than big city b with a population of over a million, it’s very important to know whether we are dealing with a cardinal measure function (it would be a very bad sign for small town a to have a higher absolute number of covid cases than big city b) or a proportional measure function (proportionally, it would not be inconceivable for small town a to have a relatively higher rate of covid cases than big city b). moreover, comparatives and superlatives about covid cases and about vaccinated people differ semantically in several theoretically interesting ways, allowing us to start to investigate how much these properties guide people’s interpretations, when coupled with different linguistic forms. the results provide new evidence that, in certain contexts, comparatives and superlatives can refer to scales ranging over degrees of proportion, in addition to degrees of cardinality. furthermore, the experimental data points to a cardinality bias (a preference for cardinal interpretations), but also shows that this bias can be weakened in favor of proportional readings by semantic factors and by certain linguistic forms. more specifically, overall, there are more cardinal interpretations with comparatives and superlatives about covid cases than with those about vaccinated people, which may be related to the former being a stage-level property and the latter being an individual-level property. these patterns are also modulated by lexical semantics (e.g. words such as rate that refer to proportions), and by whether the sentence includes an explicit mention of cases/people, which seems to make cardinal readings more available. many questions remain open for future work. in particular, a more systematic disentangling of different semantic factors is necessary in order to understand what aspects of the differences between statements about covid cases and statements about vaccinated people is responsible for the asymmetries revealed by the experiments reported here. the current work contributes a necessary foundation by providing novel evidence that these two kinds of statements clearly do differ in terms of whether they tend to receive cardinal or proportional readings, and suggests some possible explanations for this, but future work is needed to tease apart which factor(s) are driving these effects. references bale, alan & bernhard schwarz. 2020. proportional readings of many and few: the case for an underspecified measure function. linguistics and philosophy 43. 673-699. bale, alan & david barner. 2009. the interpretation of functional heads: using comparatives to explore the mass/count distinction. journal of semantics 26(3). 217-252. proceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 152 https://doi.org/10.3765/elm https://www.elm-conference.net/ beck, sigrid. 2011. comparison constructions. in claudia maienborn, klaus von heusinger & paul portner (eds.), semantics: an international handbook of natural language meaning, 1341-1389. berlin: de gruyter mouton. carlson, gregory. 1977. reference to kinds in english. ph. d. thesis, university of massachusetts. cresswell, max j. 1977. the semantics of degree. in barbara partee (ed.), montague grammar, 261-292. new york: academic press. herburger, elena. 1997. focus and weak noun phrases. natural language semantics 5. 53-78. huettner, alison. 1984. semantics seminar paper on few and many. ms., university of massachusetts. milsark, gary. 1977. toward an explanation of certain peculiarities of the existential construction in english. linguistic analysis 3. 1-29. nunberg, geoffrey. 1995. transfers of meaning. journal of semantics 12. 109-132. partee, barbara. 1989. many quantifiers. in joyce powers & kenneth de jong (eds.), proceedings of the 5th eastern states conference on linguistics (escol), 383-402. columbus, oh: department of linguistics, ohio. romero, maribel. 2015. the conservativity of many. in thimas brochhagen, floris roelofsen, & nadine theiler (eds.), proceedings of the 20th amsterdam colloquium, 20-29. amsterdam: illc. solt, stephanie. 2018. proportional comparatives and relative scales. proceedings of sinn und bedeutung, 21(2). 1123-1140. von stechow, arnim. 1984. comparing semantic theories of comparison. journal of semantics 3. 1-77 wellwood, alexis. 2015. on the semantics of comparison across categories. linguistics & philosophy, 38(1). 67-101. westerståhl, dag. 1985. logical constants in quantifier languages. linguistics and philosophy 8. 387-413. proceedings of elm 2: 142-153, 2023 elsi kaiser: comparative ambiguities and the covid pandemic. 153 https://doi.org/10.3765/elm https://www.elm-conference.net/ “liz can buy a croissant or donut. that’s both together, right?” distinguishing target free choice from non-target modal conjunction in child french antoine cochard, angeliek van hout & hamida demirdache∗ abstract. several acquisition studies have reported that children draw free choice inferences at adult-like rates from modal disjunctive statements. this study explores an alternative explanation for children’s seemingly adult-like behavior: modal conjunction, which shares verifying and falsifying conditions with free choice. however, existing experimental setups were not able to distinguish the two. with our novel design, we were able to set apart modal conjunctive interpreters from genuine free choice interpreters, using a new type of condition: a mutually exclusive context. the results revealed that free choice inferences are not so early acquired as previously thought. in contrast to the earlier studies, only half of the children between 4 and 6 were genuine adult-like free choice interpreters. the other children either show the basic inclusive interpretation of disjunction, or, as hypothesized, a modal conjunctive interpretation. keywords. acquisition; disjunction; free choice; modal conjunction; mutually exclusive context 1. introduction. plain disjunctive statements as in (1) typically yield an exclusivity inference (1a) generated by strengthening of the basic logical inclusive meaning of disjunction in (1b). thus, (1) is commonly taken to mean that exactly one of the two disjuncts is true, that is liz either bought the croissant or the donut. it has been reported in several acquisition studies that children perform non-adult-like with such statements in that they do not generate the exclusivity inference (1a) and allow the logical inclusive meaning of disjunction (1b) (chierchia et al. 2001, 2004; a.o.), or/and derive an incorrect conjunctive inference (1c) (braine & rumain 1981, cochard et al. 2023, paris 1973, singh et al. 2016, tieu et al. 2017). (1) liz bought a croissant or a donut. ‘c or d’ a. exclusivity inference: liz bought a croissant or a donut, but not both. (𝐶 ∨ 𝐷) ∧ ¬(𝐶 ∧ 𝐷) b. logical inclusive meaning: liz bought a croissant or a donut, possibly both. (𝐶 ∨ 𝐷) c. conjunctive inference: liz bought a croissant and a donut. (𝐶 ∧ 𝐷) when disjunction is embedded under a possibility modal, henceforth modal disjunctive statements, as in (2), it is inferred that both disjuncts must each be individually permitted. this type of inference has been referred to as a free choice inference (kamp 1973). so, (2) is taken to mean that there are two pastries that liz can buy, namely a croissant and a donut (2a). on top of the free choice inference, an additional, although optional (simons 2005), exclusivity inference can be ∗for helpful discussion and feedback, we would like to thank reviewers & audiences at elm3, the lling and the clcg acquisition lab. authors: antoine cochard, university of groningen & nantes university (antoinecochard@outlook.com); angeliek van hout, university of groningen (a.m.h.van.hout@rug.nl); hamida demirdache, nantes university (hamida.demirdache@univ-nantes.fr). proceedings of elm 3: 104-116, 2025 c©2025 antoine cochard, angeliek van hout, and hamida demirdache published by the lsa with permission of the author(s) under a cc by license. 104 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ derived according to which it is not possible for liz to buy the croissant and the donut, (2b). (2) liz can buy a croissant or a donut. ‘can c or d’ a. free choice inferences: i. { liz can buy a croissant. ^𝐶 } ^𝐶 ∧ ^𝐷ii. { liz can buy a donut. ^𝐷 b. exclusivity inference: { liz can buy a croissant or a donut, but not both together. ^(𝐶 ∨ 𝐷) ∧ ¬^(𝐶 ∧ 𝐷) in contrast to children’s non-adult-like interpretation of plain disjunctive statements, recent studies focusing on modal disjunctive statements suggest that children correctly derive the free choice inference at adult-like rates (huang & crain 2020, tieu et al. 2016, zhou et al. 2013). 2. disentangling free choice from modal conjunction. experiments that investigated the acquisition of free choice inferences did so by falsifying one of the two disjuncts. in this section, we show that this is not enough to confidently assert that children know how to derive such inferences. to be more precise, we investigated the idea that children’s allegedly adult-like responses might also arise in a non-adult-like grammar where they posit modal conjunction, as paraphrased in (3). (3) liz can buy a croissant and a donut. ^(𝐶 ∧ 𝐷) as a starting point, take (4) adapting the experiment setup from huang & crain (2020). in the control condition (4a), the character is permitted to do both actions: the target response is to accept the test sentence (4) since the free choice inference is satisfied. to the contrary, in the critical condition (4b), the character is only permitted to perform one of the two actions: the target response is to reject the test sentence, since the free choice inferences are no longer both satisfied. (4) liz can buy a croissant or a donut. a. control condition: liz has the possibility to buy a croissant. ^𝐶 } ^𝐶 ∧ ^𝐷she also has the possibility to buy a donut. ^𝐷 b. critical condition: liz has the possibility to buy a croissant. ^𝐶 } ^𝐶 ∧ ¬^𝐷but she does not have the possibility to buy a donut. ¬^𝐷 however, the conclusion that accepting the test sentence in (4a) and rejecting it in (4b) necessarily reflects a free choice inference is too fast. indeed, this answer pattern may also arise for another reason than correctly deriving a free choice inference: if the test sentence is incorrectly interpreted as a modal conjunctive statement, on a par with (3), then the same response pattern is expected, as illustrated in (5) for the control condition and (6) for the critical condition. (5) expected responses for the control condition (4a): true if understood as ^𝐶 ∧ ^𝐷 free choice true if understood as ^(𝐶 ∧ 𝐷) modal conjunction (6) expected responses for the critical condition (4b): false if understood as ^𝐶 ∧ ^𝐷 free choice false if understood as ^(𝐶 ∧ 𝐷) modal conjunction proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 105 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ nevertheless, a modal conjunctive interpretation is not equivalent to a free choice interpretation. this is because modal conjunction entails free choice, but not conversely. under the experimental setup presented in (4), they do indeed both have the same satisfying and falsifying conditions, but this is not always the case. a different experiment setup – namely, mutually exclusive scenarios such as (7) – allows us to tease apart their different truth conditions. (7) mutually exclusive scenario: a. i. liz has the possibility to buy a croissant. ^𝐶 } ^𝐶 ∧ ^𝐷ii. she also has the possibility to buy a donut. ^𝐷 b. liz does not have the possibility to buy both pastries. ¬^(𝐶 ∧ 𝐷) in this situation, the character has the possibility of buying either the croissant or the donut: he is free to buy whichever he chooses, as stated by (7a). but, critically, the character does not have the possibility to buy the croissant and the donut, as stated in (7b). in this scenario, free choice is satisfied (8a), but modal conjunction is not (8b). (8) liz can buy a croissant or a donut. mutually exclusive scenario a. true if understood as ^𝐶 ∧ ^𝐷 free choice b. false if understood as ^(𝐶 ∧ 𝐷) modal conjunction given that some children tend to assign a conjunctive interpretation for simple disjunctive statements, it is plausible that they may also assign a modal conjunctive interpretation to modal disjunctive statements. importantly, previous studies investigating the acquisition of free choice inferences did not include any sort of mutually exclusive condition. thus, it is premature to conclude that 4 to 6-year-olds actually derive genuine free choice inferences at adult-like rates. instead, some of the allegedly adult-like children might have given an adult-like response that arises for a non-adult-like reason. addressing this issue, we developed a new experimental design distinguishing free choice inferences from modal conjunctive ones. 3. a novel setup to investigate free choice inferences. this section presents a novel design taking inspiration from liu (2017) to investigate the derivation of free choice inferences. this design introduces two new conditions not used before in acquisition studies to our knowledge: a mutually exclusive condition adapted that allows to detect modal conjunctive interpreters and a package condition that investigates whether the exclusivity inference is computed. 3.1. participants. 68 native french children participated in the experiment. to be included in the study, participants had to be native french speakers and not have known cognitive, visual or hearing impairments. in addition, 40 native french adults were recruited on prolific1. 3.2. procedure. a truth-value judgment task (crain & thornton 1998) was used for this study. participants had to judge a sentence describing the picture of a shop-selling situation as in figure 1. each participant was told that the shop employee (the character at the bottom of figure 1) tries to help customers (e.g. the character at the top) by telling them what they can buy with the number of coins in their purse. participants had to decide whether or not the shop employee gave 1the present experiment received approval from the cedis ethics committee of nantes université (iorg0011023) and was assigned the identification number 06112023-1. proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 106 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: sample item for liz can buy a croissant or a donut. a successful description of the picture by pressing a thumbs-up button (correct) or a thumbs-down button (incorrect). the children were tested individually in a quiet space at their school using a tablet. the adults participated in a web-based version of the experiment. 3.3. stimuli. each picture contains several key elements. the bottom character, the koala, is a shop employee and appears from one item to the next; he utters the test sentence. participants were told that the shop employee is a newcomer and, as such, he can make mistakes since it is his first day at work. the top character is a customer, who varies across pictures. each customer comes to the shop with a certain number of coins in their purse. for instance, figure 1 depicts a condition where the bear arrives in the shop with one coin in her purse. finally, the shop stand shows two spaces, each of them containing a food item and under it the price. here, the croissant costs one coin and the donut costs two coins. after hearing the test sentence, the thumb-up and down buttons appear on the right side of the screen. the position of those buttons was counterbalanced between participants. 3.4. design. the experiment included the three critical conditions in figure 2 and the two control conditions in figure 3. the conditions varied the price of the food and the number of coins in the purse. it was a within-subject design and participants saw six items of each condition (18 critical items and 12 control items, 30 items in total). this was needed to determine the consistency in individual answer patterns comparing a given participant’s interpretation across the three critical conditions. (a) free choice false (b) mutually exclusive (c) package figure 2: critical conditions proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 107 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (a) control true (b) control false figure 3: control conditions for each condition, participants heard a modal disjunctive statement preceded by a lead-in fragment introducing the number of coins in the purse as in (9). (9) avec with une une pièce, coin liz liz l’ours the.bear peut can.3sg.prs acheter buy.inf un a croissant croissant ou or un a donut. donut ‘with une coin, liz the bear can buy a croissant or a donut.’ the free choice false condition was designed in the same spirit as previous studies: one of the disjuncts is accessible while the other is not. in figure 2a, with one coin, it is true that the customer can buy a donut, but it is false that she can buy a croissant, as the latter costs two coins. thus, the free choice inferences do not hold here. moreover, the modal conjunctive interpretation does not hold either, for the same reason, and also because the customer cannot buy the croissant and the donut together. so, rejection of this condition reflects either a free choice interpretation or a modal conjunctive interpretation. furthermore, a yes-answer in this condition is a clear signal of a non-adult-like interpretation. but at this stage it is not yet possible to distinguish an inclusive interpretation, where at least one disjunct is permitted, from an exclusive interpretation, where exactly one disjunct is permitted. the mutually exclusive condition is the critical condition which, alongside the free choice false condition, was needed to disentangle a free choice interpretation from a modal conjunctive one. in this condition, the possibility of buying both food items is excluded, falsifying modal conjunction. crucially, it allows a free choice interpretation, as it is possible to buy one or the other pastry, whichever she chooses. in figure 2b, with one coin, it is true that the customer can buy a croissant and it is also true that she can buy a donut. but, it is false that she can buy both the donut and the croissant as she does not have enough coins. only on the basis of an answer pattern where the participant gives a consistent yes-answer to the mutually exclusive condition alongside a consistent no-answer to the free choice false condition, can it be confidently concluded that the free choice inference is derived. the display of the pastries and price in the food stand in the package condition differs from the other two conditions. here, both food items must be bought together in a package. in figure 2c, the package containing both the donut and the croissant costs one coin. it is not possible to get the croissant without getting the donut, and conversely. this condition serves two purposes: on the one hand, it checks whether the optional exclusivity inference shown in (2b) is derived. if so, then rejection is expected. combined with the two other conditions, the package condition is needed to distinguish two potential adult-like grammars that both yield genuine free choice inferences but differ in the derivation of the exclusivity inference. on the other hand, although not critical in distinguishing a genuine free choice interpretation from a modal conjunctive one, it allows the latter and gives a chance to a modal conjunctive interpreter to give a positive answer in one of the critical proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 108 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ conditions. table 1 below summarises the expected responses patterns for each of five possible grammars. free choice false mutually exclusive package ta rg et + free choice + exclusivity (^𝑃 ∧ ^𝑄) ∧ ¬^(𝑃 ∧𝑄) no yes no + free choice − exclusivity (^𝑃 ∧ ^𝑄) ∧ ^(𝑃 ∧𝑄) no yes yes n on -t ar ge t modal conjunction ^(𝑃 ∧𝑄) no no yes − free choice − exclusivity ^(𝑃 ∨𝑄) yes yes yes − free choice + exclusivity ^(𝑃 ∨𝑄) ∧ ¬^(𝑃 ∧𝑄) yes yes no table 1: response patterns and their corresponding possible grammar if the conclusion from previous studies that children can derive a genuine free choice inference is correct, then it is expected that most children will follow one of the two possible target patterns depending on whether they draw the exclusivity inference: the [+ free choice + exclusive] or the [+ free choice – exclusive] response pattern. however, if our hypothesis is on the right track and children interpret modal disjunctive statements as modal conjunction instead, then a subset of children are expected to fit the modal conjunctive response pattern. besides the three critical conditions, twelve filler items served the purpose of balancing the number of yes and no-responses. two types of fillers were created for different purposes. six asked about elements of the display to make sure that participants were paying attention to the picture and to the recorded sentence, e.g. by stating the price of a given disjunct as in (10). the other six checked if participants understood the coin system, for example by stating what could be purchased with a certain amount of coins as in (11). (10) un a croissant croissant coûte cost.3sg.prs une one pièce. coin ‘a croissant costs one coin.’ (11) avec with une one pièce, coin liz liz l’ours the.bear peut can.3sg.prs acheter buy.inf un a croissant. croissant ‘with one coin, liz the bear can buy a croissant.’ 4. results. exclusion of participants from analysis was determined based on their performance on the controls and fillers. participants were not allowed to make more than one error out of six true controls, no more than one error out of six false controls and no more than three errors out of proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 109 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ twelve obvious yes/no filler items. after exclusion, 57 children (range = 3;11–6;9, m = 5;5) and 37 adults (range = 22-67, m = 34) remained for analysis. group results and mixed model analyses are reported in section 4.1 while individual analyses are reported in section 4.2. 4.1. group analyses. figure 4 shows that children performed differently from adults in the three critical conditions. to determine whether children’s responses significantly differed from adults, a generalized linear mixed-effects model was fitted in r (r core team 2022) using the glmer function from the lme4 package (bates et al. 2015). the model included the interaction2of group (adults vs. children) and condition (free choice false, mutually exclusive and package) as fixed effects. the controls conditions were not introduced in the model since individual variation is expected across conditions. participants were included as random effects with a random slope on conditions. taking adults as the reference point, this model revealed a significant main effect of the group (𝑝 < 0.0001), suggesting that children’s responses differed from the adults at the group level (table 2). figure 4: bars show the mean percentage of acceptance for each condition and each group. whiskers represent a confidence interval of 95% to further investigate all contrasts, tukey-corrected pairwise comparisons were computed (table 3). the comparisons confirmed that children’s responses to the free choice false condition differed from those of adults (𝑝 < 0.0001). as for the mutually exclusive condition, although no difference was found (𝑝 = 0.76), the ceiling acceptance in the adult results is problematic as it prevents us from confidently reaching the conclusion that children performed at adult-like rates for this condition.3furthermore, the comparisons also confirmed that children and adults differed significantly in their acceptance rate of the package condition, with children being more likely to accept this condition than adults (𝑝 < 0.0001). 2an anova for linear models showed that an interaction model was more successful in explaining the variance than an additive model (𝜒2 (4) = 56.5, 𝑝 < 0.0001). items were not included as random effects since each item features a unique combination of character and food. a model including items as random effects returns little to no variance at all for this factor; which means that no particular items stand out and bias the results. 3indeed, as broadly discussed in bates et al. (2015), and more specifically for linguistic data in clark et al. (2023), proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 110 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ## formula: answer ˜ condition * group + (1 + condition | part_number) ## fixed effects: ## estimate std. error z value pr(>|z|) ## fc_false:adults (intercept) -7.317 1.227 -5.963 0.00000000247 *** ## mutually_exclusive 28.027 61.202 0.458 0.64700 ## package 4.727 1.571 3.010 0.00262 ** ## children 5.063 1.232 4.109 0.00003977355 *** ## mutually_exclusive:children -23.217 61.204 -0.379 0.70443 ## package:children 2.236 2.134 1.048 0.29472 table 2: mixed model output with adults and children ## formula: pairwise ˜ group * condition, adjust = "tukey" ## $‘simple contrasts for group‘ ## adults children ## estimate se df z.ratio p.value ## condition = fc_false: -5.06 1.23 inf -4.109 <.0001 *** ## condition = mutually_exclusive: 18.15 61.21 inf 0.297 0.7668 ## condition = package: -7.30 1.82 inf -4.001 0.0001 *** table 3: pairwise comparisons output finally, in order to check whether age was a driving factor in explaining children’s responses, a generalized linear mixed-effects model was computed including only the children. this model included the critical conditions and the age in months as fixed effects. similarly to the previous model, participants were included as random effects with a random slope on conditions. this model revealed that age played no role (𝑝 = 0.495), table 4. 4.2. individual results. to answer our research questions, it is necessary to run an analysis of the individual results to determine which interpretive patterns were at stake. to do so, we defined the following criterion: if a participant consistently accepted a condition at least 5 times out of 6, she is considered as an accepter. if the condition was accepted 0 or 1 time, she was considered as a rejecter. any other number was considered to be chance behavior. table 5 shows the distribution of participants after applying this criterion. for the adults, the individual analysis matched the overall results: all except one rejected the free choice false condition and all accepted the mutually exclusive condition. however, while they agreed on the first two conditions, the group split into two uneven groups for the package condition, the majority of them rejecting it and about a quarter accepting it. in contrast, for the children, the individual analysis showed a more diverse distribution. indeed, although the majority in the free choice false and the mutually exclusive conditions resembled the adult cohort, there were some children who consistently accepted the former and/or rejected the latter. in the package condition, the children also split into two groups but in a different manner than adults: while most model convergence and overfitting issues can be aggravated by specific data scenarios such as ceiling data as it becomes nearly impossible for the model to estimate the effects of predictors accurately. to mitigate the issue of ceiling data, an option would be to compute a separate penalized regression using bayesian priors for instance. we propose such models in cochard et al. (2024). proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 111 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ## formula: answer ˜ condition + age_in_months + (1 + condition | part_number) ## fixed effects: ## estimate std. error z value pr(>|z|) ## fc_false (intercept) -1.4576 0.2857 -5.101 0.000000337 *** ## mutually_exclusive 3.2327 0.2570 12.580 <.000000001 *** ## package 3.7126 0.2727 13.616 <.000000001 *** ## age_in_months -0.1711 0.2509 -0.682 0.495 table 4: mixed model output within children free choice false mutually exclusive package adults accepter 0 37 9 rejecter 36 0 25 at chance 1 0 3 children accepter 11 41 42 rejecter 38 11 5 at chance 8 5 10 table 5: distribution of individual patterns across conditions adults consistently rejected this condition, the majority of the children accepted it and only a few of them consistently rejected it. the next step was to determine combined patterns by checking the conditions together, focusing on the free choice false and the mutually exclusive conditions. tables 6 and 7 show the distribution of response patterns for adults and children,4respectively. unsurprisingly, all adults fell into the free choice interpretation category (bottom left cell of table 6). to the contrary, children were divided across three distinct but consistent patterns: 22 children fell into the same category as adults. 11 children accepted both the free choice false and the mutually exclusive conditions, corresponding to a [– free choice – exclusive] interpretation, and 11 other children rejected both conditions, revealing a modal conjunctive interpretation. 5. general discussion. the goal of this study was to investigate how children interpret modal disjunctive statements in french, putting to test the earlier claims that children derive the associated free choice inference at adult-like rates. specifically, we tested the hypothesis that alleged adult-like children might instead be modal conjunctive interpreters, and, furthermore, investigated whether children draw the additional exclusivity inference. to carry out these goals, we developed an innovative design that included two new conditions: the first condition is the mutually exclusive condition which allows us to tease apart an (adult-like) free choice interpretation from a (non-adult4to preserve space, at-chance participants in at least one of the two conditions do not appear on tables 6 and 7. proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 112 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ‘liz can buy a croissant or a donut.’ ^(𝐶 ∨ 𝐷) mutually exclusive accept reject free choice false accept 0 inclusive ^(𝐶 ∨ 𝐷) 0 logically impossible reject 36 free choice ^𝐶 ∧ ^𝐷 0 modal conjunction ^(𝐶 ∧ 𝐷) table 6: distribution of the number of adult individual response patterns for the free choice false and the mutually exclusive conditions (n = 36/37)5. ‘liz can buy a croissant or a donut.’ ^(𝐶 ∨ 𝐷) mutually exclusive accept reject free choice false accept 11 inclusive ^(𝐶 ∨ 𝐷) 0 logically impossible reject 22 free choice ^𝐶 ∧ ^𝐷 11 modal conjunction ^(𝐶 ∧ 𝐷) table 7: distribution of the number of child individual response patterns for the free choice false and the mutually exclusive conditions (n = 44/57). like) modal conjunction interpretation, which previous experimental designs failed to differentiate. the results for this condition validate our hypothesis by showing that some children, who appear adult-like in rejecting the free choice false condition, are in fact modal conjunctive interpreters. using mutually exclusive scenarios thus proved essential in teasing apart two subgroups who agreed on the free choice false condition but differed on the mutually exclusive condition, successfully identifying a subset of children who gave a modal conjunctive interpretation of modal disjunctive statements. the second condition, the package condition, allows us to further test for the exclusivity inference, be it on the target (free choice), or the non-target (modal conjunctive) interpretations of modal disjunction. going back to table 1, four out of the five the expected response patterns were attested, summarized in table 8 below.6 proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 113 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ children adults target [+ free choice] [+ exclusivity] 4/44 (9%) 25/36 (69%) [– exclusivity] 16/44 (36%) 9/36 (25%) non-target modal conjunction 11/44 (25%) 0/36 (0%) [– free choice – exclusivity] 11/44 (25%) 0/36 (0%) table 8: summary of the four attested patterns starting with the adult-like patterns, our results show that half of consistent 4 to 6-year-olds (22/44; 50%) computed a free choice inference in an adult-like way, successfully rejecting the free choice false condition and accepting the mutually exclusive condition. however, while this subset corresponds to the largest group of children in our study, it does not reach the adult-like rates observed in previous studies where the majority of children appeared to derive a free choice inference (e.g. 98% of rejection of a condition that falsifies free choice inferences in huang & crain 2020). furthermore, adding the package condition to the discussion highlights a stark contrast between adults and children: while most of the adults rejected the package condition and derived an exclusivity inference (25/34; 73%), only 4 [+ free choice] children (4/44; 9%) did so. to the contrary, 16 children (16/44; 36%) accepted the package condition, not deriving the exclusivity inference. turning now to the non-adult-like patterns, our study found that children can face difficulties with modal disjunctive statements: first, 11 children (11/44; 25%) allowed a [–free choice – exclusivity] interpretation, accepting both the free choice false and the mutually exclusive conditions. these children show the basic inclusive interpretation of disjunction, drawing neither a free choice, nor an exclusivity inference in modal contexts. second, the last 11 children (11/44; 25%) allowed a modal conjunctive interpretation, rejecting both the free choice false and the mutually exclusive conditions. these results were confirmed in child romanian by bleotu et al. (2024) with an experimental setup replicating our conditions. conjunctive interpretations of disjunction have long been reported in the child language literature, going back to paris (1973), for matrix disjunction in plain, affirmative contexts (see §1),7 but also for disjunction embedded under every (singh et al. 2016). the novel finding reported here is that these non-target interpretations extend to modal free choice contexts.8 references bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67(1). https://doi.org/10.18637/j ss.v067.i01. 62 out of the 22 [+ free choice] children and 2 out of the 36 [+free choice] adults were at chance on the package condition. for a complete report of the package condition, see cochard et al. (2024). 7for recent eye tracking evidence with toddlers, see lobina et al. (2023). 8this finding is indeed expected, once we assume cochard et al. (2024) uniform analysis of children’s conjunctive interpretations of disjunction, across plain/affirmative, universally quantified and modal contexts, straightforwardly extending singh et al. (2016). proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 114 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ bleotu, adina camelia, lyn tieu, andreea nicolae, anton benz, gabriela bı̂lbı̂ie & mara panaitescu. 2024. free choice and exclusivity in child and adult romanian. presented at the ”free choice inferences: theoretical and experimental approaches” workshop. braine, martin & barbara rumain. 1981. development of comprehension of “or”: evidence for a sequence of competencies. journal of experimental child psychology 31(1). 46–70. https://doi.org/10.1016/0022-0965(81)90003-5. chierchia, gennaro, stephen crain, maria teresa guasti & rosalind thornton. 2001. ”some” and ”or”: a study on the emergence of logical form. in s. catherine howell, sarah a. fish & thea keith-lucas (eds.), proceedings of the 24th annual boston university conference on language development, 22–44. cascadilla press. sommerville, ma. chierchia, gennaro, maria teresa guasti, andrea gualmini, luisa meroni, stephen crain & francesca foppolo. 2004. semantic and pragmatic competence in children’s and adults’ comprehension of or. in ira a. noveck & dan sperber (eds.), experimental pragmatics, 283– 300. london: palgrave macmillan uk. https://doi.org/10.1057/9780230524125_13. clark, robert, wade blanchard, francis hui, ran tian & haruka woods. 2023. dealing with complete separation and quasi-complete separation in logistic regression for linguistic data. research methods in applied linguistics 2(1). 10–44. https://doi.org/10.1016/j.rm al.2023.100044. cochard, antoine, hamida demirdache & angeliek van hout. 2023. interpreting disjunction across positive and negative contexts: evidence from child french. in paris gappmayr & jackson kellog (eds.), proceedings of the 47th annual boston university conference on language development, 159–172. somerville, ma: cascadilla press. cochard, antoine, angeliek van hout & hamida demirdache. 2024. acquiring modal disjunctive statements: free choice vs. modal conjunction. [unpublished manuscript]. crain, stephen & rosalind thornton. 1998. investigations in universal grammar: a guide to experiments on the acquisition of syntax and semantics. cambridge, ma: mit press. huang, haiquan & stephen crain. 2020. when or is assigned a conjunctive inference in child language. language acquisition 27(1). 74–97. https://doi.org/10.1080/10489223.2 019.1659273. kamp, hans. 1973. free choice permission. proceedings of the aristotelian society 74. 57–74. publisher: aristotelian society, wiley. liu, ying. 2017. interpreting disjunction under deontic modals: an experimental investigation: utrecht university ma thesis. lobina, david j., hermann bulf & maria teresa guasti. 2023. the comprehension of conjunction and disjunction by toddlers. https://doi.org/10.17605/osf.io/s7k26. [osf preregistration]. paris, scott g. 1973. comprehension of language connectives and propositional logical relationships. journal of experimental child psychology 16(2). 278–291. https://doi.org/10.1 016/0022-0965(73)90167-7. r core team. 2022. r: a language and environment for statistical computing. r foundation for statistical computing. https://www.r-project.org/. simons, mandy. 2005. dividing things up: the semantics of or and the modal/or interaction. proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 115 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ natural language semantics 13(3). 271–316. https://doi.org/10.1007/s11050-004 -2900-7. singh, raj, ken wexler, andrea astle-rahim, deepthi kamawar & danny fox. 2016. children interpret disjunction as conjunction: consequences for theories of implicature and child development. natural language semantics 24(4). 305–352. https://doi.org/10.1007/ s11050-016-9126-3. tieu, lyn, jacopo romoli, peng zhou & stephen crain. 2016. children’s knowledge of free choice inferences and scalar implicatures. journal of semantics 33(2). 269–298. https: //doi.org/10.1093/jos/ffv001. tieu, lyn, kazuko yatsushiro, alexandre cremers, jacopo romoli, uli sauerland & emmanuel chemla. 2017. on the role of alternatives in the acquisition of simple and complex disjunctions in french and japanese. journal of semantics 34(1). 127–152. https://doi.org/10.109 3/jos/ffw010. zhou, peng, jacopo romoli & stephen crain. 2013. children’s knowledge of free choice inferences. in todd snider (ed.), salt 23, 632–651. cornell university. proceedings of elm 3: 104-116, 2025 antoine cochard, angeliek van hout, and hamida demirdache: distinguishing target free choice from non-target modal conjunction in child french. 116 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ on a concessive reading of the rise-fall-rise contour: contextual and semantic factors alexander göbel & michael wagner* abstract. this paper presents three auditory rating experiments on the rise-fall-rise contour (rfr). experiment 1 provides experimental evidence that the rfr makes disagreeing with a prior statement more natural than neutral intonation would. additionally, the data show that the rfr exhibits a valence asymmetry, noted by göbel (2019): the amelioration of a disagreement is greater when the rfr is used in a positive reply to a negative statement than in a negative reply to a positive statement. experiments 2 and 3 investigate factors contributing to this asymmetry, showing that it disappears in replies to questions and is weakened when the reply contains an additive particle. based on these results, we argue that the rfr has a scalar meaning, following göbel (2019), with the relevant scale being contextually determined and resulting in an ambiguity resembling the focus-particle at least. keywords. intonation; focus-particles; alternatives 1. introduction. this paper addresses the question how intonation affects and conveys linguistic meaning through the lens of intonational contours, or “tunes”, specifically, with a case-study on the so-called rise-fall-rise contour (rfr). the rfr has been characterized as conveying a sense of uncertainty or incompleteness on the speaker’s part (ward & hirschberg 1985).1 in apparent opposition to this characterization, the rfr can also be used more antagonistically, as in the naturally occuring example in (1). here, dewey’s reply seems to challenge hal’s prior statement, but without straightforwardly contradicting it. (1) hal: i mean, you can’t just break into a zoo, role a couple of elevens and suddenly become ... the dean of a university. dewey: i did... (malcolm in the middle: s7, e20; audio) the goal of this paper is to substantiate the tension between uses of the rfr as in (1) and its prior characterization empirically and explore relevant factors by presenting three auditory rating studies. we argue that there are in fact distinct uses of the rfr that resemble an ambiguity of the english focus-particle at least, and support the proposal by göbel (2019) that the rfr indicates the presence of a higher alternative on a variable scale. the structure of the paper is as follows. section 2 discusses details of prior accounts of the rfr and introduces the idea of a parallel between the rfr and at least. section 3 presents the three experiments and section 4 provides discussion of the results. 2. background. *we want to thank massimo lipari and emma nguyen for providing recordings for the studies, andrea beltrama, chris potts, maribel romero, and the audience at elm2 for feedback on the project, as well as a feodor-lynen fellowship of the humboldt-foundation to the first author and an serc discovery grant to the second author for funding. authors: alexander göbel, mcgill university (alexander.gobel@mcgill.ca) & michael wagner, mcgill university (chael@mcgill.ca). 1we will focus here on the meaning contribution of the rfr and leave equally important prosodic issues aside. proceedings of elm 2: 83-94, 2023 c©2023 alexander göbel and michael wagner published by the lsa with permission of the author(s) under a cc by license. 83 https://doi.org/10.3765/elm https://www.elm-conference.net/ 2.1. prior accounts. the seminal account by ward & hirschberg (1985) is primarily concerned with cases as in (2) where the rfr intuitively is used as a hedge that avoids a more definitive answer to the question while still intending to be relevant. ward & hirschberg’s proposal centers the use of the rfr around speaker uncertainty in relation to a scale and its scalar values, either regarding the appropriateness of evoking such a scale, the particular choice of scale, or the choice of a newly added value on a given scale. applied to (2), we could then say that there is uncertainty about “the relative positions of missouri and the mississippi on a geographical scale ordered from east to west” (w&h: 767). (2) a: have you ever been west of the mississippi? b: i’ve been to missouri... (ward & hirschberg 1985, (62)) in more recent work, the intuition that the rfr conveys uncertainty has been implemented in terms of incompleteness by constant (2012) and wagner (2012). for constant, the rfr resembles focus-particles like only in that it quantifies over alternatives and conveys that all assertable alternatives (= those alternatives whose truth-value is not yet known) cannot be safely claimed, formalized as in (3) as a conventional implicature.2 wagner proposes a similar analysis, with the difference that (i) relevant alternatives are speech acts rather than propositions and (ii) that there is an alternative speech act that could have been made, rather than focusing on what alternatives are not claimable, see (4).3 (3) jrfr ϕkci = ∀p ∈ jϕkf s.t. p is assertable in c: the speaker cannot safely claim p. (4) jrfrk = λs. ∃s’ in jskga, s ↛ s’ and performing s’ might be justified: s these two accounts capture the distribution of the rfr in quantifier scales as in (5), where the rfr is felicitous with the middle value some but infelicitous with the extreme values none or all. none is odd because it rules out all higher alternatives such that the rfr applies vacuously since there are no alternatives left to be quantified over (constant) or there is no alternative speech act to be made (wagner), and all is odd for similar reasons since another alternative such as some is already entailed by it. in contrast, some leaves alternatives open that can then be marked as not being claimable or as the alternative speech act that was not made, hence the characterization of the rfr as conveying incompleteness (see also goodhue et al. 2016 for experimental evidence). (5) a: did you feed the cats? b: i fed { #none / some / #all } of them... (with rfr) a different approach comes from westera (2019) (see also westera 2017), who focuses on the role of the rfr in the discourse structure, specifically the question under discussion (qud, roberts 2012). according to westera, the rfr indicates that a conversational maxim in relation to the main qud has been suspended and instead a secondary qud is being addressed. a potential benefit of this view is that it may allow including cases of contrastive topic, that share a similar prosodic profile, into a unified theory. however, to what extent such a unification is warranted is an open 2constant (2012) also discusses the role of the rfr in scope ambiguities like all...not, which we put aside here, but see syrett et al. (2014) for experimental data on this issue. 3another difference regards the association with focus, which is obligatory in constant’s theory but optional in wagner’s. we will come back to this issue in the general discussion. proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 84 https://doi.org/10.3765/elm https://www.elm-conference.net/ question. 2.2. göbel (2019). the main approach that this paper builds upon comes from göbel (2019). the central data point provided there is the observation that the rfr shows an asymmetry in replies to statements that contrast in the valence attributed to what is being discussed: while the rfr is natural as a response to a’s negative characterization of dexter in (6a) when descriptively opposing the previous statement with a positive characterization, the reverse a negative reply to a positive statement is markedly odd. we will dub this contrast in acceptability the “valence asymmetry”. (6) valence asymmetry a. a: dexter is such a horrible person. b: he gives to charity... (audio) b. b: dexter is such a great person. a: ?? he murders people... (audio) notably, this data point is unexpected on the accounts discussed in the previous subsection. to evaluate ward & hirschberg (1985), let’s assume that the scale under consideration is a “goodness” scale, and what the replies convey is uncertainty about where to locate dexter in virtue of his actions. it is unclear why (6a) would be more appropriate in this case than (8b) given that knowing someone gives to charity is compatible with being uncertain whether they are a horrible person, and knowing that someone murders people is compatible with being uncertain whether they are a great person.4 for constant (2012) and wagner (2012), it is equally unclear why (6a) should leave alternatives open but not (6b), and for westera (2019) there is no obvious difference in terms of the secondary quds (6a) and (6b) give rise to either. the solution göbel (2019) proposes is that the rfr indicates that there is an alternative that ranks higher on a given scale, essentially borrowing from ideas present in prior accounts. let’s again imagine a goodness scale that goes from pure evil on one end to pure virtue on the other, illustrated in figure 1. figure 1: illustration of scalar effects for (6). in (6a), a’s statement places dexter toward the bottom of the scale. b’s reply then suggests that dexter is located higher. in contrast, in (6b) b’s statement starts dexter off toward the top and a’s reply adds a negative counterpoint. as a consequence, the carrier utterance of the rfr does not contribute a higher alternative but a lower one and results in oddness.5 the account additionally 4admittedly, ward & hirschberg’s account allows further possible interpretations of where the issue lies, but it is that flexibility that renders it less explanatory. 5a potential objection, however, would be that (6) is not a fair comparison: giving to charity is a less good indicator that someone is not a horrible person than murdering people is for indicating that someone is not a great person. one goal of the following experiment is to provide quantitative data that allows us to confirm or deny the reliability of the valence asymmetry. proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 85 https://doi.org/10.3765/elm https://www.elm-conference.net/ predicts the contrast in (5): both none and all fail to leave higher alternatives open, while some leaves all.6 2.3. parallel to at least. interestingly, the valence asymmetry in (6) is not unique to the rfr, but can also be observed with concessive (or evaluative) at least (nakanishi & rullmann 2009). the characteristic feature of concessive at least is that alternatives are ranked according to desirability, for instance in (7) conveying that feeding some of the cats is still better than feeding none of them. (7) at least cameron fed some of the cats (... it could’ve been less). when using concessive at least in the same dialogues as (6) but with a falling contour, we see the same intuitively even clearer pattern: at least is felicitous with a positive reply (8a), while it is infelicitous with a negative reply (8b) (8) a. a: dexter is such a horrible person. b: at least he gives to charity. b. b: dexter is such a great person. a: # at least he murders people. the parallel of the rfr with at least crucially does not end there. in addition to the concessive interpretation, at least also has an epistemic interpretation that has been tied to uncertainty (geurts & nouwen 2007). taking (9) as illustration, the speaker only commits to cameron having fed no less than some of the cats while leaving open the possibility that she fed all of them. (9) cameron fed at least some of the cats (... maybe even all of them). this meaning of epistemic at least resembles prior accounts of the rfr that focus on its uncertainty or incompleteness component.7 the resemblance goes as far as epistemic at least showing the same acceptability pattern with quantifiers as we saw in (5), namely being incompatible with extreme ends of a scale: (10) a: did cameron feed the cats? b: she fed at least { #none / some / #all } of them. based on this parallelism of the rfr with distinct interpretations of at least, we propose that the rfr is prone to a similar type of ambiguity. this view would allow us to reconcile theoretical accounts that emphasize the uncertainty and incompleteness component of the rfr with the valence asymmetry. moreover, it would allow us to make sense of cases like (11), where the rfr is felicitous despite the truth of the higher alternative ten being known, showing that uncertainty or incompleteness are relevant but not sufficient. on the ambiguity view of the rfr, (11) would simply be a case of a concessive usage. 6note that the case in (6) and the one in (5) differ on this view with respect to where the relevant alternatives come from: in (6), the alternative comes from the host utterance in relation to a previous one present in the discourse, while in (5), it is alternatives to the element that receives focus that are relevant. we will come back to this issue later. 7a notable tension of this view comes from experimental results reported in de marneffe & tonhauser (2019), which show that the rfr strengthens implicatures when compared to neutral intonation. this finding is at odds with likening the rfr to an implicature suspension device like epistemic at least. proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 86 https://doi.org/10.3765/elm https://www.elm-conference.net/ (11) a: did you feed all ten cats? b: i didn’t feed all ten, but i fed nine of them... (with rfr) (audio) however, an initial obstacle to this approach is the question what factors contribute to the ambiguity and determine the respective interpretation. we will explore this question in experiments 2 and 3. beforehand, experiment 1 will be dedicated to substantiating the valence asymmetry empirically. 3. experiments. 3.1. experiment 1. 3.1.1. materials & design. the main goal of this experiment was to provide quantitative data on the valence asymmetry. to do so, we used short dialogues modeled after (6) that varied the “valence” of each utterance, shown in (12). (12) a. negative + positive (= “mismatch”) a: the bike ride yesterday was really terrible, the weather was horrific. b: we had a cocktail... (neutral, rfr) b. positive + negative (= “mismatch”) a: the bike ride yesterday was really great, the weather was perfect. b: we had an accident... (neutral, rfr) c. negative + negative (= “match”) a: the bike ride yesterday was really terrible, the weather was horrific. b: we had an accident... d. positive + positive (= “match”) a: the bike ride yesterday was really great, the weather was perfect. b: we had a cocktail... the cases in (12a)-(12b) are labeled “mismatch” conditions since the valence of the context utterance and the target utterance differ. to complete the paradigm, we also added cases where valences where the same as a type of control, labeled “match” conditions.8 additionally, each target utterance was recorded in a neutral falling contour as a baseline and with an rfr. the design was thus a 2x2x2 (context-valence: negative vs positive, response match: match vs mismatch, intonation: neutral vs rfr). we created 8 item sets like (12). each item set was presented in all eight conditions but ordered in a way so that each participant saw each item set in all different conditions before seeing the item set again (within-design), for a total of 48 item trials. 3.1.2. procedure. the experiment was implemented through prosodyexperimenter (https: //github.com/prosodylab/prosodylabexperimenter) and ran online on prolific.ac. participants first saw a welcome screen, followed by a chance to adjust their volume and test their microphone, an online consent form, and a language background questionnaire. afterwards, there was a test where participants were played three sounds and had to choose which one was the quietest, which required the use of headphones. for the main part of the experiment, participants were 8the manipulation could also be viewed in terms of whether speaker b agrees or disagrees with a’s statement. we avoid using these labels here, however, since the extent to which a speaker’s utterance is interpreted as a (dis)agreement might be affected by intonation, such that focusing on the propositional content is more neutral. proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 87 https://doi.org/10.3765/elm https://www.elm-conference.net/ auditorily presented with a dialogue and had to provide a naturalness rating on a scale from 1 (completely unnatural) to 6 (completely natural) based on how they thought the response sounded given the context. there were three practice trials after receiving instructions, followed by 48 stimuli. the experiment concluded with a chance to provide feedback. a test version of the experiment can be accessed at https://prosodylab.org/˜agobel/conepi/conaddratingbare/ ?session_id=test&mode=experiment. 3.1.3. participants. 29 participants were recruited from prolific.ac and paid $2.20 each. five participants were excluded due to failing the headphone check, leaving 24 for data analysis. 3.1.4. results. coding & data analysis. context-valence was sum coded given there was no basis for choosing one over the other as a baseline. intonation and response match, on the other hand, were dummy coded, with match and neutral as reference levels. data were modeled with ordinal mixed effects with the maximal structure allowing convergence. results and analysis files with more details on the full models of this and the following experiments, as well as experimental files and stimuli, can be found at the associated osf repository: https: //osf.io/t36qk the average ratings by condition are shown in figure 2. the most notable although less interesting pattern concerns higher ratings for match than mismatch conditions, reflected in a significant effect of response match (z = 17.19, p < .001***). looking at only match conditions next, we see higher ratings for neutral intonation compared to the rfr, reflected in a significant effect of intonation (z = -7.20, p < .001***). however, when comparing the intonation difference for match with mismatch, the penalty for the rfr decreases, reflected in a significant interaction between intonation and response match (z = 5.30, p < .001***). moreover, this amelioration of the penalty is greater with negative context sentences and hence positive target sentences than with positive context sentences, evidenced by a three-way interaction of intonation, response match, and context-valence (z = 2.25, p < .05*). figure 2: naturalness ratings by condition, experiment 1. proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 88 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.1.5. discussion. the experiment revealed several notable findings. first, participants strongly preferred dialogues where context utterance and target utterance had the same valence over dialogues that featured some opposition. while this was the numerically largest effect, it is also of least interest to this investigation such that we will put it aside. second, the rfr ameliorated this mismatch penalty. albeit novel, this finding is in line with prior accounts of the rfr insofar as expressing uncertainty or incompleteness may be viewed as more polite in an otherwise confrontational dialogue and hence improve ratings. finally and most crucially, the extent to which the rfr ameliorated the penalty was larger in dialogues where the context utterance was negative and the carrier utterance of the rfr contributed a positive counterpoint compared to dialogues with a positive starting statement and a negative counterpoint. this last pattern is evidence for the valence asymmetry observed in göbel (2019) and the argument made in this paper, namely that there is a concessive reading of the rfr alike to at least. a follow-up question to this argument, mentioned in section 2.3, is what makes this reading of the rfr available. a relevant hint in this regard comes from the fact that the prior literature has mostly focused on uses of the rfr in replies to questions or out of context, whereas göbel (2019) and experiment 1 investigated replies to (value) statements. a possible contrast between questions and statements is furthermore in line with intuitions and may thus be expected to contribute to the choice between an “epistemic” and a “concessive” reading of the rfr. the next experiment examines this connection. 3.2. experiment 2. 3.2.1. materials & design. the goal of this experiment was to test whether replies to questions exhibit the same valence asymmetry we observed for replies to statements in experiment 1. all design aspects of this experiment were identical to experiment 1, with the sole exception that context utterances were changed from statements to questions, as in (13): (13) a. negative + positive (= “mismatch”) a: do you think yesterday’s bike ride was terrible? b: we had a cocktail... b. positive + negative (= “mismatch”) a: do you think yesterday’s bike ride was great? b: we had an accident... c. negative + negative (= “match”) a: do you think yesterday’s bike ride was terrible? b: we had an accident... d. positive + positive (= “match”) a: do you think yesterday’s bike ride was great? b: we had a cocktail... 3.2.2. procedure. the procedure was identical to experiment 1. a test version of the experiment can be found here: https://prosodylab.org/˜agobel/conepi/conaddratingq/ ?session_id=test&mode=experiment 3.2.3. participants. 30 participants were recruited from prolific.ac and paid $2.20 each. six participants were excluded due to failing the headphone check, leaving 24 for data analysis. proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 89 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.2.4. results. coding & data analysis were the same as experiment 1 (see section 3.1.4). the average ratings by condition are shown in figure 3. the first thing to note is that we again observe a mismatch penalty, with higher ratings for match than for mismatch, reflected in a significant response match effect (z = -8.89, p < .001***). additionally, ratings for positive contexts are lower than negative contexts in match cases, reflected in a significant effect of context-valence (z = 4.81, p < .001***), but higher in mismatch cases, which shows up as a significant interaction of context-valence and response match (z = -5.47, p < .001***). the second notable pattern is that there is little to no difference between neutral intonation and rfr across the board, with the exception of the match conditions in positive contexts, resulting in a marginal effect of intonation (z = 1.82, p < .1•), a significant interaction of intonation and context-valence (z = -2.13, p < .05*), a significant interaction of intonation and response match (z = -1.96, p < .05*), and a marginal three-way interaction between intonation, context-valence, and response match (z = 1.65, p < .1•). figure 3: naturalness ratings by condition, experiment 2. 3.2.5. discussion. as in experiment 1, participants considered dialogues that matched in valence more natural than those that mismatched. moreover, this mismatch penalty was larger when the context question was negative than when it was positive. however, since this effect was largely independent of intonation, we will put it aside here. more crucially, using questions instead of statements as in experiment 1 seemed to render the difference between neutral intonation and the rfr almost non-existent. the only exception were higher ratings for the rfr in the match condition in positive contexts. most importantly, there was no evidence for a valence asymmetry like we observed in experiment 1. while the relevant three-way interaction was marginally significant, it went in the opposite direction to what we would have expected, namely with the rfr increasing the mismatch penalty in positive contexts rather than decreasing it less. the results thus provide support for the idea that using the rfr to reply to a question or to a statement affects what reading it receives. proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 90 https://doi.org/10.3765/elm https://www.elm-conference.net/ following göbel (2019), we assume that the two readings can be modeled in terms of different scales that are determined pragmatically, for instance mediated by a qud (e.g. beaver & clark 2008). for statements as in experiment 1, the scale is evaluative, as discussed in sec 2.2. for questions as in this experiment, the relevant notion is truth, which can be modeled in terms of a scale of informativity. what is at stake in the question whether the bike ride was terrible or great then is to what extent the reply provides a definitive answer. crucially, both positive and negative replies leave open “stronger”, more definitive answers, such that the rfr is equally acceptable independently of valence, in contrast to what we saw in experiment 1. the following experiment explores a potential semantic factor for the interpretation of the rfr, namely the role of additive particles, based on the intuition that adding such a particle to the rfr utterance weakens the valence asymmetry. 3.3. experiment 3. 3.3.1. materials & design. the goal of this experiment was to test whether an additive particle like also would neutralize or at least weaken the valence asymmetry observed in experiment 1. design and materials were identical to experiment 1, except that the target utterances now contained also, as in (14): (14) a. negative + positive (= “mismatch”) a: the bike ride yesterday was really terrible, the weather was horrific. b: we also had a cocktail... (neutral, rfr) b. positive + negative (= “mismatch”) a: the bike ride yesterday was really great, the weather was perfect. b: we also had an accident... (neutral, rfr) c. negative + negative (= “match”) a: the bike ride yesterday was really terrible, the weather was horrific. b: we also had an accident... d. positive + positive (= “match”) a: the bike ride yesterday was really great, the weather was perfect. b: we also had a cocktail... 3.3.2. procedure. the procedure was identical to experiments 1 and 2. a test version of the experiment can be accessed here: https://prosodylab.org/˜agobel/conepi/ conaddrating/?session_id=test&mode=experiment. 3.3.3. participants. we recruited 30 participants from prolific.ac, who received $1.60 each. six participants were excluded due to failing the headphone check, leaving 24 for data analysis. 3.3.4. results. coding & data analysis were the same as experiments 1 and 2 (see section 3.1.4). the average ratings by condition are shown in figure 4. the results pattern largely resembles that of experiment 1: the biggest effect is a penalty for mismatch conditions compared to match conditions (z = -15.13, p < .001***), neutral intonation is rated better than rfr in match conditions (z = -5.26, p < .001***), but rfr makes mismatch cases less bad than neutral intonation (z = 5.05, p < .001***). however, in contrast to experiment 1, the extent to which the mismatch penalty proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 91 https://doi.org/10.3765/elm https://www.elm-conference.net/ amelioration of the rfr is greater in negative contexts than positive ones is less clear here, and in fact not supported by a significant three-way interaction (z = 0.63, p = .53). figure 4: naturalness ratings by condition, experiment 3. 3.3.5. discussion. the results provided some support for the intuition that an additive particle weakens the valence asymmetry of the rfr: even though both the mismatch penalty and the amelioration of it by the rfr were comparable to experiment 1 where also was absent, there was no statistical evidence that the amelioration differed between the two contexts. however, the numerical trend was still in line with a valence asymmetry. moreover, the relevant comparison here is between the presence of an effect in experiment 1 and the absence of the same effect in this experiment, rather than based on an effect within the same experiment, and should thus be taken as tentative evidence only. a proper confirmation that additive particles weaken the valence asymmetry will therefore be left for future research. we may nonetheless wonder why additive particles would have such an effect. the potential explanation we want to put forward here is that the way the reading of the rfr gets determined interacts with the meaning of also. additive particles are commonly analyzed as presupposing the truth of an alternative proposition (e.g. heim 1992). in doing so, the interpretation of the rfr may become biased toward an “epistemic” reading that is concerned with informativity where valence is irrelevant and hence the valence asymmetry gets weakened. this approach would point toward an interesting issue of compositionality between two types of meanings focus-particles and intonational contours that have been treated including here as making similar contributions based on relations to alternatives. how to spell out this connection in detail, however, goes beyond the scope of this paper and will be left for future work. 4. general discussion. this paper presented data from three auditory rating experiments on the rfr. experiment 1 substantiated a contrast between the acceptability of the rfr in positive replies to negative statements and negative replies to positive statements, as argued by göbel (2019), proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 92 https://doi.org/10.3765/elm https://www.elm-conference.net/ which we dubbed the valence asymmetry. we interpreted this finding as evidence for a “concessive” reading of the rfr, resembling concessive at least. experiment 2 investigated replies to questions rather than statements and showed that the valence asymmetry disappears in this case. we took this change to point toward questions inducing a distinct, “epistemic” reading, again borrowing from a parallel with at least. on this view, the two readings can be unified by assuming that the rfr indicates the presence of a higher alternative but on different scales that are pragmatically determined. finally, experiment 3 explored the intuition that the inclusion of an additive particle like also weakens the valence asymmetry, for which we found tentative evidence. if confirmed, this finding raises interesting questions about how the meaning of intonational contours interacts compositionally with different parts of its carrier utterance. apart from this issue, there are two others we want to briefly discuss here. the first concerns the role of alternatives in the advocated theory. as mentioned earlier, the two readings are not taken to only differ in the type of scale involved but also in how they generate alternatives. for the “concessive” reading from experiment 1, the relevant higher alternative is contributed by the host utterance of the rfr, with a prior statement being the proposition that the alternative is higher than. for the “epistemic” reading from experiment 2, the higher alternative is an alternative to the host utterance itself, similar to how the assumed alternative to some in the quantifier cases (5) from section 2.1 is all. a prediction of this view then is that the rfr should in principle be compatible with extreme scale values as long as it contributes an alternative that is higher than some previously mentioned one and receives a concessive reading.9 future research will have to show whether this prediction is borne out. the second issue is about the meaning of intonation at large. here, we took the rfr to contribute its meaning holistically rather than attempting to attribute associate smaller meaning components with the individuals parts of the contour, but we can ask what such a decompositional account may look like. the rfr is prosodically analyzed as, using tobi annotation, consisting of an l*+h pitch accent, a lphrasal accent, and a h% boundary tone. we will put the phrasal accent aside here and focus on the pitch accent and the boundary tone. according to pierrehumbert & hirschberg (1990), the l*+h accent evokes a scalar meaning, which is consistent with the view of the rfr adopted here. regarding the high boundary tone, there are a number of recent accounts of rising declaratives to draw from. rudin (2022) proposes that the characteristic feature of rising intonation is that they lack speaker commitment, which seems incompatible with the valence asymmetry given that for both types of replies the speaker is intuitively committed to the truth of the proposition. jeong (2018) argues that rising intonation opens a meta-linguistic issue, which again seems to lack a clear path for explaining the valence asymmetry per se, but may capture the intuitive sense of indirectness and the resulting politeness effect. we aim to address the question about the contribution of the final rise in future work. references beaver, david & brady clark. 2008. sense and sensitivity: how focus determines meaning. oxford: blackwell. 9this difference is furthermore reminiscent of and possibly connected to the issue between constant (2012) and wagner (2012) regarding whether the rfr obligatorily associates with focus or not. proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 93 https://doi.org/10.3765/elm https://www.elm-conference.net/ constant, noah. 2012. english rise-fall-rise: a study in the semantics and pragmatics of intonation. linguistics and philosophy 35. 407–442. geurts, bart & rick nouwen. 2007. ‘at least’ et al.: the semantics of scalar modifiers. language 83. 533–559. https://doi.org/10.1353/lan.2007.0115. goodhue, daniel, lyana harrison, y. t. clémentine su & michael wagner. 2016. toward a bestiary of english intonational contours. in brandon prickett & christopher hammerly (eds.), proceedings of the north east linguistics society 46, 311–320. göbel, alexander. 2019. additives pitching in: l*+h signals ordered focus alternatives. proceedings of semantics and linguistic theory (salt) xxix 1–11. heim, irene. 1992. presupposition projection and the semantics of attitude verbs. journal of semantics 9. 183–221. jeong, sunwoo. 2018. intonation and sentence-type conventions: two types of rising declaratives. journal of semantics 35. 305–356. de marneffe, marie-catherine & judith tonhauser. 2019. inferring meaning from indirect answers to polar questions: the contribution of the rise-fall-rise contour. in edgar onea, malte zimmermann & klaus von heusinger (eds.), questions in discourse, 132–163. leiden: brill. nakanishi, kimiko & hotze rullmann. 2009. epistemic and concessive interpretation of at least. talk presented at canadian linguistics association. https://linguistics.sites. olt.ubc.ca/files/2018/03/2009.nakanishi_rullmann.cla_-1.pdf. pierrehumbert, janet b. & julia hirschberg. 1990. the meaning of intonational contours in the interpretation of discourse. in p. r. cohen, j. morgan & m. e. pollack (eds.), intensions in communication, 271–311. cambridge, ma: mit press. roberts, craige. 2012. information structure in discourse: towards an integrated formal theory of pragmatics. semantics and pragmatics 5. 1–69. https://doi.org/10.3765/sp.5.6. earlier version appeared in osu working papers in linguistics 49 in 1996. rudin, deniz. 2022. intonational commitments. journal of semantics 39. 339–383. syrett, kristen, georgia simon & kirsten nisula. 2014. prosodic disambiguation of scopally ambiguous quantificational sentences in a discourse context. journal of linguistics 50. 453– 493. wagner, michael. 2012. contrastive topics decomposed. semantics and pragmatics 5. 1–54. ward, gregory & julia hirschberg. 1985. implicating uncertainty: the pragmatics of fall-rise intonation. language 61. 747–776. westera, matthijs. 2017. exhaustivity and intonation: a unified theory: university of amsterdam dissertation. westera, matthijs. 2019. rise-fall-rise as a marker of secondary quds. in daniel gutzmann & katharina turgay (eds.), secondary content: the linguistics of side issues, leiden: brill. proceedings of elm 2: 83-94, 2023 alexander göbel and michael wagner: on a concessive reading of the rise-fall-rise contour. 94 https://doi.org/10.3765/elm https://www.elm-conference.net/ disambiguating quantity judgements: mass/count and extra-grammatical cues sven smeman, maaike smit, james a. hampton, & yoad winter* abstract. comparative quantity judgements are a useful probe into the semantics of the mass/count distinction, where count nouns usually trigger cardinal comparisons (more dogs), and mass nouns trigger non-cardinal measurement (more rice). however, exceptions like ‘object’ mass nouns (furniture) and ‘mixed’ comparatives (more gold than diamonds) complicate this pattern. in such cases there is often a mismatch between the mass/count status of the noun and the criterion for comparison, which challenges our understanding of the mass/count distinction and how it affects quantity judgements. we propose that these mismatches reflect a systematic ambiguity, where the mass/count distinction is one of the factors influencing disambiguation. using a new experimental method focused on ambiguity judgements instead of truth-value judgements, the results support the traditional semantic encoding of the mass/count distinction, with operations of ‘packaging’ and ‘grinding’ triggered by extra-grammatical factors. keywords. mass/count nouns; comparatives; disambiguation; grinding; packaging 1. introduction. the mass/count distinction in english is often viewed as a grammatical reflection of semantic discreteness, with count nouns (e.g. cats) typically referring to discrete objects and mass nouns (e.g. water, justice) representing substances or abstract concepts without a clear atomic structure. however, english and many other languages have ‘object’ nouns like furniture and jewelry that challenge this correlation, as their referents are discrete despite being grammatically mass. recent analyses suggest weaker links between grammatical countability and semantic discreteness, either by downplaying the role of grammar with object mass nouns (gillon 1992) or proposing discrete interpretations of such nouns (barner & snedeker 2005, chierchia 2010). further research has shown that quantity judgements also rely on pragmatic pressures and the perceived discreteness of real-world referents (scontras et al. 2017, rothstein 2017), which are not always clear-cut, leaving the semantics of the mass/count distinction an ongoing area of debate. this paper studies potential ambiguities in quantity judgements with comparatives. instead of focusing on categorical judgments of cardinality or measurement as in barner & snedeker’s work, we explored participants’ perception of ambiguities between these strategies. results show that ambiguity in comparatives is closely tied to the mass/count distinction, both with ‘flexible’ nouns (e.g. more rocks/rock) and with ‘object’ mass nouns (e.g. more bags/baggage). however, simple noun expressions (bags, baggage, rock/s) do not support the same levels of unambiguous counting or measurement that is seen with expressions like number of bags/rocks or volume of baggage/rock. these results reflect the gradient of perceived ambiguity shown in figure 1. the ambiguity of quantity judgments has important implications for theories linking the mass/ count distinction to continuous/discrete meanings. while our results show systematic evidence *work on this paper was supported by a grant of the european research council (erc) under the european union’s horizon 2020 research and innovation programme (grant agreement no. 742204). sven smeman, maaike smit, & yoad winter: utrecht university, james a. hampton: city st. george’s, university of london. proceedings of elm 3: 359-370, 2025 c©2025 sven smeman, maaike smit, james a. hampton, and yoad winter published by the lsa with permission of the author(s) under a cc by license. 359 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: scale of ambiguity in quantity judgements for this connection, it is just one factor influencing counting and measurement. the behavior of ‘object’ mass nouns, previously seen as exceptional, aligns with their conventionally discrete referents. however, even ‘substance’ mass nouns like beer can exhibit counting effects in contexts that highlight packaging, though less so than object mass nouns. likewise, the strong tendency for count nouns to trigger cardinal comparisons can be contextually inhibited. we argue that these observations support a theory that applies ‘grinding’ and ‘packaging’ operations based on pragmatic triggering. in this theory, grinding and packaging are not only triggered by mismatches between syntax and lexical semantics as in too much rabbit (‘grinding’ of a count noun) or three beers (‘packaging’ of a mass noun), but also without such mismatches as in more rabbits (or more beer), in case pragmatics triggers grinding (packaging, resp.), leading to exceptional quantity judgements. 2. background: mass/count distinctions and quantity comparisons. in examples like (1), count nouns (cns, e.g. bag/s) are in the plural and support a cardinal comparison: (1) anna has more bags/clay/baggage than ben. by contrast, mass nouns (mns) are in the singular and show an irregular behavior: ‘substance’ mns (smns, e.g. clay) prefer non-cardinal measurement (by volume, weight, etc.) while ‘object’ mns (omns, e.g. baggage, furniture) trigger cardinal comparisons (mccawley 1975). barner & snedeker (2005) experimentally studied quantity judgments by presenting participants with comparative questions on visual stimuli as in figure 2. the stimuli depicted two quantities, one of which with fewer discrete items whose total size is larger. this creates a scenario where counting and measurement should yield different answers to the question who has more ⟨noun⟩?. the omns in this study all showed the same (near unanimous) levels of counting as the cns. figure 2: who has more shoes/toothpaste/silverware? (barner & snedeker 2005) figure 3: who has more? (scontras et al. 2017) the special properties of omns have led researchers to relax the strict correspondence between the mass/count distinction and continuity/discreteness of noun meanings. gillon (1992) suggested that while all cns have discrete, countable meanings, the grammar of mns is ‘mute’ on discreteness. thus, mns are expected to support discrete/continuous interpretations only on the basis of extra-grammatical factors (e.g. world knowledge). this line does not offer any generalproceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 360 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ization on the semantics of mns. by contrast, barner & snedeker introduce a sharp distinction between two types of mns: smns vs. omns. omns are likened to cns using a syntactic feature (+ind = ‘individual’) that distinguishes both noun classes from smns (−ind).1 a problem for this account comes from examples like (2) below (rothstein 2017, grimm & levin 2012): (2) john has more (#pieces of) furniture than bill, so he should use the larger moving truck. rothstein argues that omns like furniture in (2) can be measured by volume, unlike cns like tables, which must be counted. rothstein generalizes that cns are counted in comparatives, smns are measured, and omns can be either counted or measured.2 this view aligns with gillon’s, attributing the smn/omn distinction to extra-grammatical factors. scontras et al. (2017) experimentally showed the importance of such factors by observing that non-cardinal measurement often occurs with objects perceived as substances, even if not explicitly referred to as mns (figure 3). hampton & winter (2024) further support rothstein’s claims experimentally, demonstrating a greater reliance on non-cardinal measurement for omns like baggage compared to cns like bags in contexts favoring such comparisons. 3. ambiguity in quantity comparisons. what is the general role that the mass/count distinction and extra-grammatical factors play in the quantity interpretation of nouns? this question involves two empirical problems: cn problem: to what extent can cns compromise count-based interpretations? mn problem: to what extent can mns compromise measure-based interpretations? specifically, do omns support counting as strongly as cns? there is general agreement that omns frequently allow cardinality-based interpretations in comparatives despite their syntactic ‘mass’ type. however, following hampton & winter (2024), we question whether omns reject measure-based interpretations to the same extent as cns. this prompts further investigation into ‘exceptional’ quantity interpretations of both mns and cns, as illustrated in the following examples: (3) anna ate more beans/peas/lentils than ben. (4) mary put two (and a half) oranges in the punch. (5) a. pirates’ treasures usually contained more gold than diamonds. b. he had more hair than teeth, and his hairs totalled three. in (3), we are more likely to refer to amounts of beans, peas and lentils than to their cardinalities (mccawley 1975). sentence (4) predominantly reports on the cardinality of oranges, but at least for some speakers, it also involves a ‘container’ reading, which refers to amounts of orange juice (snyder 2021). comparatives that ‘mix’ cns and mns as in (5) may lead to a non-cardinal measurement of diamonds as in (5-a), and cardinality judgements about hair as in (5-b) (winter 1chierchia (2010) makes a similar distinction in semantic terms: smns have an ‘unstable’ atomic structure whereas omns have atoms that are as ‘stable’ as those of cns. thus, like barner & snedeker, chierchia assumes a categorical distinction between nouns in terms of their ‘discreteness’, putting omns and cns in one class and smns in another. 2for rothstein, cns grammatically require counting, whereas omns support counting through measurement using numerosity estimation. proceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 361 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 2022). these cases support the idea that ‘mismatches’ between a count/mass status and a cardinality/measurement interpretation are not limited to omns. mismatches between syntactic noun classification and lexical mass/count preferences are welldocumented. for example, a typical cn like bicycle can appear as syntactically ‘mass’ (there’s bicycle all over the floor), while beer, usually an mn, can appear as countable (three beers). building on these phenomena, we propose that cns (mns) without syntactic mismatch can also invoke non-cardinal (cardinal, resp.) quantity judgments, driven by extra-grammatical or pragmatic factors. thus, phrases like a lot of bicycles (beer) or more bicycles (beer) may involve non-cardinal (cardinal, resp.) quantity judgments, making omn behavior less exceptional than thought. by weakening the semantic connection between count/mass and cardinal/non-cardinal quantity judgements, we do not intend to deny the strong effect of the mass/count distinction on interpretation. indeed, following hampton & winter (2024), we hypothesize that cns are more easily interpreted using cardinal interpretations than mns, and conversely: it is easier to apply non-cardinal measurement to mns than to cns. to summarize, we present two hypotheses: (h1) discreteness is semantically encoded in cns, whereas mns encode continuity. (h2) comparatives grammatically allow measurement and counting with both cns and mns, where the choice is affected by the mass/count distinction, together with pragmatic factors and the perception of real-world objects. to test these hypotheses we experimented with quantity interpretations, and examined to what extent they are affected by the mass/count distinction and by grammar-external factors. in our experiments, as in previous experimental work, we asked speakers to contribute linguistic judgements on comparative statements relative to situations where counting and measurement lead to different results. for example, in one of our tests we presented participants with sentences like a has more bags/baggage than b in situations where in terms of cardinality a has fewer bags, but a’s bags occupy a larger volume than b’s bags. unlike barner & snedeker (2005), we did not use forced choice questions about a single comparative sentence. rather, participants were requested to simultaneously report their judgements about two comparative sentences, of the forms: (6) a. a has more ⟨noun⟩ than b b. b has more ⟨noun⟩ than a we asked participants to indicate whether they accept both sentences in (6) or only one of them. if speakers can employ both counting and measurement, they are expected to choose the first option; if they only employ one strategy, they are expected to choose only one of the sentences. this method follows recent works that highlight problems in evaluating semantic theories using simple truth-value judgements (syrett & musolino 2015, pinto & zuckerman 2019). as syrett & musolino point out, speakers may have robust biases against one of the readings of a given sentence although their semantics does not disallow it. when sentences like (6-a) and (6-b) are presented in isolation, truth-value judgements may put a pressure on speakers to reject one of these sentences just because their salient strategy makes the other sentence true. for instance, in relation to figure 2 speakers may reject a sentence like a has more shoes than b because they strongly prefer the cardinality-based interpretation of cns, hence b has more shoes than a would be far preferred as their description of the situation. presenting both sentences simultaneously aims to reduce the effect of this potential confound. proceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 362 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ packages-post bags-baggage instruments-equipment (pieces of ) furniture weapons-weaponry stationery (items) figure 4: visual stimuli in experiment 1 4. testing ambiguity in quantity comparisons. this section reports four experiments that study the hypothesized ambiguity of quantity judgements with cns and mns. these experiments introduced participants to situations where quantity judgements are expected to vary depending on whether the comparison is cardinal or non-cardinal. each situation was accompanied by two comparative sentences as in (6) above. participants were asked to select one of these sentences, or both of them, as a possible description of the given situation. the experiments involved referents of cns and mns whose common perception is as substances (e.g. rock, clay) and discrete objects (bags, apples). the experiments tested hypotheses (h1)-(h2) about the influence of the count/mass distinction on the choice between cardinal and non-cardinal quantity judgements. specifically: • experiment 1 tested exceptional strategies of non-cardinal measurement, comparing omns (baggage) to cns (bags). • experiment 2 tested the effect of a cn/mn environment on exceptional counting strategies with smns (clay). • experiment 3 tested non-cardinal measurement with simple cns (apples) by comparing them to cn phrases that include an overt cardinal expression (number of apples). • experiment 4 tested counting with simple smns (rock) by comparing them to smn phrases that include an overt non-cardinal expression (volume of rock). these experiments are described below in more detail. 4.1. experiment 1. the aim of this experiment was to examine if, and to what extent, noncardinal measurement is tolerated with omns more easily than with coreferential cns, although the referents are preferably perceived as discrete. 4.1.1. materials and procedure. we selected six omn-cn pairs from (hampton & winter 2024): (7) packages-post, bags-baggage, instruments-equipment, sofas-furniture, weapons-weaponry, stationery items-stationery as in (hampton & winter 2024), these nouns were selected using two criteria: (i) minimal referential differences between the omn and the cn in each pair; (ii) the ease of representing different items, where some items are of a greater volume. for each noun in (7), each participant was presented with a corresponding drawing (figure 4) and two comparative sentences containing the corresponding omn (or cn). for example, with respect to the relevant drawing in figure 4, some participants were presented with the pair of sentences in (8) and others were presented with (9): (8) a. anna has more packages than ben. b. ben has more packages than anna. proceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 363 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (9) a. anna has more post than ben. b. ben has more post than anna. using sentences and (8-a) and (9-a) to describe the respective situation in figure 4 reflects counting, since anna has a larger cardinality of packages. conversely, (8-b) and (9-b) reflect noncardinal measurement. in relation to such pairs, participants were given the following question: (10) “which of the following better describes your reaction to these sentences?” subsequently, participants were asked to choose one of the following two statements: (11) a. “i imagine either one of the two sentences might be used to describe the situation” b. “only one of the sentences can be used felicitously” we interpret answer (11-a) as reflecting a judgement about the ambiguity of the given sentence between counting and measurement. answer (11-b) is interpreted as reflecting lack of ambiguity. participants who selected option (11-b) received a followup question: (12) “which of the two sentences do you think can be used to describe the image?” accordingly, the answer of each participant was coded as unambiguous counting, unambiguous measurement, or ambiguity. using prolific, we recruited 481 speakers of british english (309 female, age m=42.1). each participant received exactly one of the 12 target questions. in total, between 40-41 responses were collected for each of the 12 items. the experiment started with a ‘training’ stage. participants were introduced to the following pairs of sentences concerning a text reporting that anna’s bag weighs 20 kilos, and ben’s bag weighs 10 kilos: (13) a. anna’s bag is heavier than ben’s bag. b. ben’s bag is lighter than anna’s bag. (14) a. anna’s bag is heavier than ben’s bag. b. ben’s bag is heavier than anna’s bag. answer (11-a) fits (13) and (11-b) fits (14). participants who answered differently were given an explanation and were asked to correct their answer. this aim of this ‘training’ was to demonstrate that two sentences with argument reversal as in (8) and (9) may or may not both be true, hence to dissuade participants from adopting automatic strategies when considering such examples. 4.1.2. results. table 1 gives the total numbers of unambiguous counting, unambiguous measurement and ambiguity for cns and omns. to test (h1) we are interested in comparing noncardinal measurement with omns and cns. for this comparison, we encoded all ‘ambiguous’ judgements and unambiguous ‘measurement’ judgements as reflecting the property +measure. in our results, omns showed a +measure behavior in 46% of the cases, and cns in 16% of the cases, yielding a significant effect according to fisher’s exact test (p < 0.00001). the odds ratio was 0.23 (95% confidence interval [0.15, 0.35]). all omns showed a higher frequency of +measure judgements than their corresponding cn, with a significant effect in four of the six pairs. the +measure behavior with omns was recorded with levels between 28% and 65%, while with cns it ranged between 5% and 38%, with only one cn (stationery items) showing +measure in more than 20% of the cases. the per item effects are summarized in figure 5. 4.1.3. discussion. the results strengthen the claims in (grimm & levin 2012, rothstein 2017) and (hampton & winter 2024), according to which omns have a privileged access to non-cardinal proceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 364 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ experiment 1 experiment 2 experiment 3 experiment 4 cn omn cn-smn smn-smn cn no. of cn smn vol. of smn counting 202 130 64 27 54 121 27 8 measurement 14 43 55 103 63 29 103 141 ambiguity 25 67 40 31 42 12 31 10 total 241 240 159 161 159 162 161 159 table 1: unambiguous counting/measurement and ambiguity in experiments 1-4 legend: [+] : p < 0.05 [∗] : p < 0.01 [∗∗] : p < 0.001 [∗∗∗∗] : p < 0.00001 figure 5: experiment 1 – 95% confidence intervals of odds ratios, comparing noncardinal measurement in omn and cn conditions figure 6: experiment 2 – 95% confidence intervals of odds ratios, comparing counting in cnsmn and smn-smn conditions measurement compared to cns with the same referents. notwithstanding, the acceptance of the +measure strategy was not at zero levels with cns. we may account for this effect as reflecting lack of attention, which led some participants to deviate from the literal meaning of cns. however, we can also interpret it as reflecting a genuine potential of cns to trigger non-cardinal measurement. this question will be tested more directly in experiment 3. 4.2. experiment 2. the results of experiment 1 support the idea that the ‘mass’ status of omns makes their non-cardinal interpretation more accessible than with cns. omns are special among the mns in that their referents are commonly perceived as discrete objects. however, smn referents may be packaged into discrete objects like heaps of sand or puddles of water. do such situations make smns more amenable to quantity judgements based on cardinality? the aim of experiment 2 was to test this question by looking into cardinality judgements with smns, and how they are affected by the mass/count distinction in their syntactic environment. 4.2.1. materials and procedure. we selected four nouns, each of them in singular and plural: (15) chocolate-chocolates, rope-ropes, rock-rocks, stone-stones we selected these nouns because of their flexible smn/cn use, and, considering the plausibility of comparison with their referents, we selected the following smns: flour, sand, clay and soil, proceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 365 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ chocolate-chocolates (vs. flour) rope-ropes (vs. sand) rock-rocks (vs. clay) stone-stones (vs. soil) figure 7: visual stimuli in experiments 2 and 4 respectively. we used each of the 8 noun tokens in (15) with a pair of sentences like (16) and (17): (16) a. this image shows more chocolate than flour. b. this image shows more flour than chocolate. (17) a. this image shows more chocolates than flour. b. this image shows more flour than chocolates. these sentences were presented together with the corresponding drawing in figure 7. using sentences and (16-a) and (17-a) to describe the situation reflects counting, since the image shows a larger cardinality of chocolate candies. conversely, (16-b) and (17-b) reflect non-cardinal measurement. we used the same procedure as in experiment 1: a question (10) about the pair of sentences, with the two options in (11) and a followup question (12) for participants who selected option (11-b). the experiment started with the same ‘training’ stage as in experiment 1. using prolific, we recruited 320 speakers of british english (205 female, age m=41.9). each participant received exactly one of the 8 target questions. in total, between 40-41 responses were collected for each of the 8 items. 4.2.2. results. table 1 gives total numbers of unambiguous counting, unambiguous measurement and ambiguity for cn-smn and smn-smn comparatives. to test (h2), we compared cardinal judgements with smns (flour) when they appear with smns (chocolate) and cns (chocolates) as in (16)-(17), encoding all ‘ambiguous’ judgements and unambiguous ‘counting’ judgements as ‘+count’. in our results, comparisons between a cn (chocolates) and a smn (flour) showed a +count behavior in 65% of the cases. smn-smn pairs (chocolate-flour) showed a +count behavior only in 36% of the cases, which is a significant difference from cn-smn pairs (p < 0.00001) according to fisher’s exact test. the odds ratio was 0.30 (95% confidence interval [0.19, 0.47]). all cn-smn pairs showed a higher frequency of +count judgements than their smn-smn counterparts, with a significant effect in three of the four cases. the +count behavior with cn-smn pairs was recorded with levels between 53% to 74%, while with smn-smn pairs it ranged between 24% and 53%. the per item effects are summarized in figure 6. 4.2.3. discussion. the results clearly support hypothesis (h1) in terms of showing the effect of the cn/mn distinction. mixed comparisons with cns (chocolates) and smns (flour) led to significantly more counting compared to their uniform smn-smn counterparts (chocolate-flour). effects of the cn/mn distinction on quantity comparisons with ‘flexible’ nouns also appeared robustly in barner & snedeker’s experiments. in addition, our results highlight the possibility of ‘mismatches’ in quantity comparisons: mixed cn-smn comparatives showed high levels of counting with smns (65% of the participants) and non-cardinal measurement with cns (60%). however, although ‘mixed’ cn-smn environments promoted counting with smns, also smn-smn proceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 366 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ comparisons showed a considerable proportion of the participants (36%) who accepted counting: either unambiguously (17%), or in addition to non-cardinal measurement (19%). this shows that when the context makes counting salient (e.g. using heaps of flour as in figure 7), counting of smns is possible even without syntactic pressure. the tolerance of smns to counting and of cns to non-cardinal measurement is studied further in experiments 3 and 4. 4.3. experiment 3. in experiments 1 and 2, cns showed non-zero levels of tolerance towards non-cardinal measurement (16% and 60%, resp.). such effects may be grammatically licensed, but they could also reflect a misinterpretation of the utterance that deviates from its literal meaning. the aim of experiment 3 was to test to what extent cns like apples are distinguished in this respect from expressions like number of apples, where cardinality is explicitly mentioned. if the literal meanings of both expressions trigger counting to the same degree, we expect no difference between them. if the literal meanings of cns are tolerant towards non-cardinal measurement as in hypothesis (h2), we expect differences to show up. 4.3.1. materials and procedure. we selected four pairs of cns: (18) apples-almonds, bananas-hazelnuts, cod fillets-peas, potatoes-olives these pairs of cns were selected because of the different sizes of the objects they refer to, and due to their common appearance in recipes (e.g. an apple almond pie). each of these noun pairs was used with a pair of sentences as in (19) and (20) below: (19) a. anna needs more apples than almonds. b. anna needs more almonds than apples. (20) a. anna needs a greater number of apples than almonds. b. anna needs a greater number of almonds than apples. these sentences were presented together with the following scenario: (21) “for baking an apple almond pie, anna needs: 900 grams of apples: about 10 apples and 70 grams of almonds: about 50 almonds” other pairs of nouns in (18) were studied in a similar way. we used textual representations, as describing recipes using graphical stimuli proved hard. in the situation (21), using sentences (19-a) and (20-a) reflects non-cardinal measurement, as anna needs a larger weight of apples. conversely, (19-b) and (20-b) reflect counting. if measurement is prohibited with cns, we expect sentences (19-a) and (20-a) to be rejected at similar levels. we used the same procedure as in experiments 1-2: the same ‘training’ stage, then a question (10) about the pair of target sentences, with the two options in (11) and a followup question (12) for participants who selected option (11-b). using prolific, we recruited 321 speakers of british english (222 female, age m=38.0). each participant received exactly one of the 8 target questions. in total, between 40-41 responses were collected for each of the 8 items. 4.3.2. results. table 1 reports total numbers of unambiguous counting, unambiguous measurement and ambiguity for bare cns and number of cns. to test (h2), we are interested in comparing the extent to which speakers apply non-cardinal measurement with cns in sentences like (19) and (20). for this comparison, we encoded all ‘ambiguous’ judgements and unambiguous proceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 367 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ‘measurement’ judgements as reflecting the property +measure. in our results bare cns showed a +measure behavior in 66% of the cases, and number of cns only in 25% of the cases, yielding a significant effect (p < 0.00001) according to fisher’s exact test. the odds ratio was 0.18 (95% confidence interval [0.11, 0.29]). all bare cns showed a higher frequency of +measure judgements than their number of counterparts, with a significant effect in all four pairs. the +measure behavior with bare cns was recorded with levels between 61% to 77%, while with number of cns it ranged between 20% and 37%. the per item effects are summarized in figure 8. legend: [+] : p < 0.05 [∗] : p < 0.01 [∗∗] : p < 0.001 [∗∗∗] : p < 0.0001 [∗∗∗∗] : p < 0.00001 figure 8: experiment 3 – 95% confidence intervals of odds ratios, comparing non-cardinal measurement in the cn and number of cn conditions figure 9: experiment 4 – 95% confidence intervals of odds ratios, comparing counting in smn and volume of smn conditions 4.3.3. discussion. acceptance of non-cardinal measurement was substantially higher with bare cns compared to number of expressions, although the latter also showed considerable tolerance towards measurement (25% of the participants). this supports the assumption that pragmatic pressures for measurement can override literal meaning even with what seems as the clearest example of a cardinal expression. however, this divergence from literal meaning cannot explain why bare cns showed higher levels of measurement. if the literal meanings of both bare cns and number of cns categorically rule out non-cardinal measurement, the substantial difference between them remains unexplained. the large difference observed between bare cns and number of cns strengthens the conclusion that the considerable levels of measurement with cns in experiments 1 and 2 are not divergences from literal meanings, but reflect a silent ‘grinding’ effect, similarly to the use of typical cns in syntactic ‘mass’ environments (e.g. too much rabbit in the chilli), and in accordance with rothstein’s (2017) analysis of measure expressions with cns (2 kilos of nuts). 4.4. experiment 4. experiment 3 compared non-cardinal measurement with cns to phrases that explicitly mention cardinality (number of cn). experiment 4 tested the opposite question: to what extent do smns show counting effects compared to phrases that explicitly mention noncardinal measurement (volume of smn)? 4.4.1. materials and procedure. we used smn-smn comparatives with the same four pairs of smns as in experiment 2: (22) chocolate-flour, rope-sand, rock-clay, stone-soil proceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 368 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ however, each of these noun pairs was now used in pairs of sentences as in (23): (23) a. this image shows a larger volume of chocolate than flour. b. this image shows a larger volume of flour than chocolate. these sentences were presented together with the corresponding drawing in figure 7, where similarly to experiment 2, sentence (23-a) reflects counting and (23-b) reflects non-cardinal measurement, the goal was to compare pairs of sentences as in (23) to the four smn-smn pairs of experiment 2, as in (16). we used the same procedure as in experiments 1-3: the same ‘training’ stage, then a question (10) about the pair of sentences, with the two options in (11) and a followup question (12) for participants who selected option (11-b). using prolific, we recruited 159 speakers of british english. together with the participants in the smn-smn condition of experiment 2, this resulted in 320 participants (222 female, age m=39.5). each of these 320 participants received exactly one of the 8 target questions with pairs of sentences as in (16) (experiment 2) and (23) (experiment 4). in total, between 39-41 responses were collected for each of the 8 items. 4.4.2. results. table 1 reports total numbers of unambiguous counting, unambiguous measurement and ambiguity for bare smns and volume of smns. to test (h2), we are interested in comparing the extent to which speakers apply counting with smns in sentences like (16) and (23). for this comparison, we encoded all ‘ambiguous’ judgements and unambiguous ‘counting’ judgements as reflecting the property +count. in our results bare smns showed a +count behavior in 36% of the cases, and volume of smns only in 11% of the cases, yielding a significant effect (p < 0.00001) according to fisher’s exact test. the odds ratio was 0.23 (95% confidence interval [0.13, 0.41]). all smns showed a higher frequency of +count judgements than their volume of smn counterpart, with a significant effect in three of the four cases. the +count behavior with bare smns was recorded with levels between 24% to 53%, while with volume of smns it ranged between 3% and 20%. the per item effects are summarized in figure 9. 4.4.3. discussion. the acceptance of cardinality comparisons was substantially higher with bare smns compared to volume of expressions, although the latter also showed some tolerance towards cardinality (11% of the participants). in sentences like (16) (experiments 2 and 4) and (23) (experiment 4), the only pressure to apply counting with smns was pragmatic (i.e. in the visual stimulus). this is similar to how non-cardinal measurement with cns was only triggered by pragmatics in experiment 3. similarly to our conclusion about cns in experiment 3, we interpret this result as indicating that despite the strong contribution of smns to the decision on non-cardinal measurement, their literal meaning licenses pragmatic triggering of a ‘packaging’ operator, similar to the syntactically triggered interpretation of examples like three beers. 5. conclusions. the experiments in this paper show that the mass/count distinction systematically affects the interpretation of comparatives, where mismatches between noun type and quantity judgements are common beyond just object mass nouns. we conclude that all mass nouns encode continuity in their meanings, while all count nouns encode discreteness. mismatches between meanings and quantity judgements are explained by covert ‘grinding’ and ‘packaging’ operators (rothstein 2017), which can be triggered by syntax or pragmatics. thus, as our experiments demonstrated, pragmatics may pressure comparatives like more nuts or more clay to yield an proceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 369 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ atypical quantity judgement, despite the lack of mismatch between syntax and lexical mass/count preferences. this fact may have been obscured by eliciting truth-value judgements in neutral contexts, but is revealed when participants are given the opportunity to contribute judgements about ambiguity. we conclude that object mass nouns like furniture can be treated as an extreme case of pragmatic mismatches in the mass domain. aggregate count nouns like beans, peas and lentils may exemplify the opposite situation: count nouns that pragmatically favor non-cardinal measurement. further research is needed on cross-linguistic variations with the mass/count encoding of these concepts, and their effects on quantity judgements. references barner, david & jesse snedeker. 2005. quantity judgments and individuation: evidence that mass nouns count. cognition 97. 41–66. https://doi.org/10.1016/j.cognition.2004.06.009. chierchia, gennaro. 2010. mass nouns, vagueness and semantic variation. synthese 174. 99–149. https://doi.org/10.1007/s11229-009-9686-6. gillon, brendan s. 1992. towards a common semantics for english count and mass nouns. linguistics and philosophy 15. 597–639. https://doi.org/10.1007/bf00628112. grimm, scott & beth levin. 2012. who has more furniture? presented at the mass/count in linguistics, philosophy and cognitive science conference. unpublished ms. hampton, james a. & yoad winter. 2024. countability and comparative judgements. https://www.phil.uu.nl/ yoad/papers/hamptonwintersub2023.pdf. to appear in proceedings of sinn und bedeutung 28. mccawley, james d. 1975. lexicography and the count-mass distinction. in annual meeting of the berkeley linguistics society, vol. 1, 314–321. https://journals.linguisticsociety.org/proceedings/index.php/bls/article/view/2335/2105. pinto, manuela & shalom zuckerman. 2019. coloring book: a new method for testing language comprehension. behavior research methods 51. 2609–28. https://doi.org/10.3758/s13428018-1114-8. rothstein, susan. 2017. semantics for counting and measuring. cambridge: cambridge university press. https://doi.org/10.1017/9780511734830. scontras, gregory, kathryn davidson, amy rose deal & sarah e. murray. 2017. who has more? the influence of linguistic form on quantity judgments. proceedings of lsa 2. 41:1–15. https://journals.linguisticsociety.org/proceedings/index.php/plsa/article/view/4097/3785. snyder, eric. 2021. counting, measuring, and the fractional cardinalities puzzle. linguistics and philosophy 44. 513–550. https://doi.org/10.1007/s10988-020-09297-5. syrett, kristen & julien musolino. 2015. all together now: disentangling semantics and pragmatics with together in child and adult language. language acquisition 23. 175–197. https://doi.org/10.1080/10489223.2015.1067319. winter, yoad. 2022. mixed comparatives and the count-to-mass mapping. in empirical issues in syntax and semantics 14, 309–338. http://www.cssp.cnrs.fr/eiss14/eiss14 winter.pdf. proceedings of elm 3: 359-370, 2025 sven smeman, maaike smit, james a. hampton, and yoad winter: disambiguating quantity judgements: mass/count and extra-grammatical cues. 370 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ contrafactives, learnability, and production david strohmaier & simon wimmer* abstract. no natural language has contrafactive attitude verbs. because factives are universal across natural languages, this means that there is a major asymmetry between contrafactives and factives. we previously hypothesised that this asymmetry arises partly because the meaning of contrafactives is significantly harder to learn than that of factives. here we test this hypothesis by using a production-oriented computational experiment that overcomes two limitations of our previous experiments. we find that our results do not support our previous hypothesis. keywords. contrafactives; factives; semantic universals; transformers; learnability 1. introduction. no natural language appears to have what holton (2017) calls a ‘contrafactive’, i.e. a morphologically atomic attitude verb that would mirror know and other factives.1,2 like know, a contrafactive would entail a belief in the content of its declarative complement. but unlike know, it would presuppose the falsity (not truth) of that complement.3 for instance, if english had a contrafactive contra, ayesha contras that beatrice is cool would entail that ayesha believes that beatrice is cool, but presuppose that it is false that beatrice is cool. that no natural language appears to have a contrafactive is one part of an asymmetry between contrafactives and factives.4 the second part is that know, and so at least one factive, appears to have counterparts in all natural languages (goddard 2010, hannon 2015). given that a contrafactive would mirror a factive, this two-part asymmetry raises the question: why do contrafactives and factives differ so greatly in their frequency? there have been several attempts to explain this asymmetry. holton (2017) suggests that contrafactives are ruled out because there are no entities suitable to serve as the semantic values of their declarative complements.5 roberts & özyildiz (2023) suggest that they are ruled out because they do not satisfy a necessary condition on presupposition triggering: that asserted content depends on *this paper reports on research supported by cambridge university press and assessment, university of cambridge. we thank the nvidia corporation for the donation of the titan x pascal gpu used in this research. we are grateful to giulia martina as well as audiences at and anonymous reviewers for elm3 and the mecore closing workshop, especially kajsa djärv, natasha korotkova, todor koev, tom roberts, and wataru uegaki. david strohmaier designed and ran the computational experiment, simon wimmer brought philosophical and linguistic discussions to bear on design and interpretation. authors: david strohmaier, department of computer science and technology, alta institute, university of cambridge (ds858@cam.ac.uk) & simon wimmer, department of philosophy, heinrich-heine-university duesseldorf (simon.wimmer@hhu.de). 1for discussion of relevant cross-linguistic evidence see rosenberg (1975), kierstead (2015), krifka (2016), holton (2017), hsiao (2017), anvari et al. (2019), sander (2020), hoeksema (2021), bochnak & hanink (2022), bossi (2022), roberts & özyildiz (2023), strohmaier & wimmer (2022, 2023), mcgregor (2024). 2given the atomicity condition, anvari et al. (2019)’s creerse is no contrafactive, despite their use of the label. 3the presupposition condition entails that adjectives like falsely or erroneously cannot give rise to a contrafactive, since falsely believe entails, but does not presuppose, the falsity of its declarative complement. 4potential candidates not yet discussed in sufficient detail to assess whether they should count as contrafactives include pitta-pitta’s widi and other examples by mcgregor (2024) as well as german’s wähnen by sander (2020). but as we previously noted, our target asymmetry would remain even if some contrafactives were found. 5for critical discussion see hyman (2017), wimmer (2019), roberts & özyildiz (2023). proceedings of elm 3: 395-410, 2025 c©2025 david strohmaier and simon wimmer published by the lsa with permission of the author(s) under a cc by license. 395 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ presupposed content. finally, in earlier work we hypothesized that contrafactives are less common than factives partly because their meaning is harder to learn than that of factives (strohmaier & wimmer 2022, 2023).6,7 our aim here is to further explore our hypothesis. we initially motivated our hypothesis via two mismatches generated by contrafactive, but not factive, attitude ascriptions (strohmaier & wimmer 2022; 300-1, 2023; 71-3): first, a mismatch between their matrix subject’s commitment to the truth, and their speaker’s commitment to the falsity, of the contrafactive’s declarative complement; second, a mismatch between the primary use of a contrafactive’s declarative complement to make assertions (and thereby commit to its truth) and the commitment to the declarative complement’s falsity incurred by using a contrafactive attitude ascription. we expected these mismatches to make it harder to learn the meaning of a contrafactive than that of a factive.8 to test our expectation, we ran two experiments with artificial neural networks designed to capture those mismatches. the results from these experiments supported our hypothesis: in both cases, the loss dropped more slowly for contrafactives than for factives. but as section 2 explains, our previous experiments had two key limitations. they did not distinguish falsity from presupposition failure. and they only tested comprehension, not production. we therefore conducted a third experiment designed to overcome those limitations. yet as section 3 shows, the results from this experiment do not support our previous hypothesis. we find no consistent difference between contrafactives and factives. for reasons we sketch in our conclusion (section 4), however, we are not yet sure whether this undermines our previous hypothesis. 2. previous experiments. in strohmaier & wimmer 2022, 2023 we used transformer encoders with binary cross entropy loss from the pytorch library. we trained them to predict the truth value of factive, non-factive, and contrafactive attitude ascriptions based on three inputs: an attitude ascription; a mind representation to tell the model what the matrix subject of the ascription believes about the world; and a world representation to tell the model what the world is like.9 we encoded the mind and world representations differently across the two experiments. in ‘experiment 1’, reported in strohmaier & wimmer 2022, they consisted of sequences of tokens that we glossed as representing a 3x3 grid filled with varying geometrical objects of varying colours (e.g. ‘blue-square empty empty ...’);10 in ‘experiment 2’, reported in strohmaier & wimmer 2023, the representations consisted of a single token representing a proposition (e.g. ‘r p110’). given this difference, we also encoded the attitude ascriptions we fed into the model differently across experiments 1 and 2. to illustrate, in experiment 1 contrafactive attitude ascriptions looked like contra blue square above red triangle; in experiment 2 like contra p110. our transformers produced the same kind of output in both experiments: a value within [0,1]. this value encodes an estimate of the probability that an input attitude ascription is true, given 6and, as steinert-threlkeld & szymanik (2019; 4) note, natural languages intuitively tend to use compositional, non-atomic means to express meanings that are harder to learn. 7maldonado et al. (2022) show that humans can learn contrafactives. partly for this reason, we did not intend our hypothesis to fully explain the asymmetry between contrafactives and factives. 8our emphasis on the first mismatch was inspired by literature on theory of mind, especially phillips & norby (2021), our emphasis on the second by work on pragmatic-syntactic bootstrapping, especially hacquard & lidz (2022). 9our paradigm for a non-factive attitude ascription is a belief ascription. like know and a contrafactive, believe entails a belief in the content of its declarative complement. however, it presupposes neither truth nor falsity. 10we did not train the model on visual information. our gloss is merely intended to ease human interpretation. proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 396 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the input mind and world representations. a value of 1 encodes an estimate of ‘definitely true’; 0 encodes one of ‘definitely not true.’ in evaluating our models’ learning performance, we considered the distance between their estimates, given the input world and mind representations, and the correct value (1 or 0), given those representations. in both experiments, the rolling mean of the loss fell more slowly for contrafactives than for factives. in experiment 1, for instance, the mean loss after 100,000 training examples for contrafactives was 0.54, for factives 0.39. our transformers approached the correct values of contrafactive attitude ascriptions more slowly than that of their factive counterparts. 2.1. limitations. we previously noted several limitations of our experiments. some hold for one experiment, but not the other. for instance, our transformer in experiment 1 may have made some mistakes partly due to difficulties transformers have with word order (compare pham et al. 2021). this was one key reason why experiment 2 used simpler inputs. a more general limitation is the question of whether transformers approximate human language learning closely enough for us to draw conclusions about humans from results about transformers. we previously referred to several encouraging results (e.g. caucheteux & king 2022, merkx & frank 2021, schrimpf et al. 2021) that show correlations between transformer and human performance. for instance, schrimpf et al. argue that transformers (especially gpt-2) explain a high proportion of all explainable variance from brain measurements (fmri and ecog) during sentence processing tasks, such as reading and listening tasks. we can add to this list of encouraging results here. kallini et al. (2024) argue that transformers learn english more easily than languages that humans cannot learn. paape (2023) suggests that transformers approximate human performance for depth charge illusions and their non-illusory counterparts. there are also encouraging results for attitude verbs specifically. as ziembicki et al. (2023) explain, transformers (bert in particular) approximate the factivity of polish attitude verbs as judged by polish-speaking expert linguists more closely than polish-speaking non-expert linguists. similarly, ross & pavlick (2019) suggest that transformers (again bert) closely approximate human judgments about the veridicality of english attitude verbs. admittedly, transformers do not match human performance perfectly. still, the overall trend is clear: transformers closely approximate human performance. given this, we continue to tentatively draw conclusions about humans from results about transformers. 2.2. motivation for current experiment. our latest experiment is motivated by, and addresses, two further limitations of both of our previous experiments. the first is that an output value of 0 only encodes an estimate that the input ascription is definitely not true. for instance, if in experiment 2 our transformer outputs 0 for contra p110, we take our transformer to estimate that contra p110 is definitely not true. but this estimate does not distinguish whether the ascription is false or a presupposition failure. in fact, our transformer cannot distinguish these cases. to see this, contrast two cases. in case 1, our transformer gets r p110 as mind and world representation. this world representation verifies the declarative complement p110. this makes the ascription a presupposition failure. in case 2, our transformer gets input mind and world representations distinct from r p110. a non-r p110 world representation falsifies the declarative complement and a non-r p110 mind representation means the matrix subject does not believe p110. contra p110’s falsity presupposition is satisfied, proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 397 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ but its belief entailment is not. this makes the ascription false. the problem is that our model treats both cases alike: in both cases, we train it to output 0. that our transformer should distinguish these cases does not depend on our view of presuppositions. of course, if presupposition failure yields undefined or a third truth value, rather than falsity, our transformers must learn to distinguish false from undefined or third value ascriptions to approximate human performance. but even if presupposition failure yields falsity, our transformers should treat ‘mere’ falsity and presupposition failure differently. for that is just what human language users do: the ‘hey! wait a minute’ diagnostic (e.g. von fintel 2004) suggests as much. the second key limitation of both of our previous experiments is that they only consider how easily our transformers learn to comprehend the attitude ascriptions we feed into them. (they successfully learn to comprehend them iff their output values for those ascriptions, given some input mind and world representations, match the correct values for those ascriptions, given those representations.)11 but this means that the results of our experiments do not speak to the relative ease of learning how to produce contrafactive and factive attitude ascriptions. our transformers got attitude ascriptions as inputs; they did not produce them as outputs. this is problematic because transformers may learn how to comprehend and to produce contrafactive attitude ascriptions at different speeds. ideally, they also learn how to produce contrafactive attitude ascriptions more slowly than to produce factives ones. this would strengthen our hypothesis. but pessimistically, transformers may also more quickly learn how to produce contrafactive attitude ascriptions than factive ones. this would raise questions about our hypothesis. if the meaning of a contrafactive is easier to learn than that of a factive in a production-oriented setting, but harder to learn in a comprehension-oriented setting, can a difference in learnability explain any of the asymmetry in frequency between contrafactives and factives? to address the two key limitations just noted, we ran a third computational experiment. now, our transformer learns how to produce the expressions whose meaning it has to learn: factive, non-factive, and contrafactive attitude ascriptions. this immediately addresses the second key limitation and yields a paradigm that, as far as we know, has not been used yet in the computational literature on learnability and semantic universals (compare steinert-threlkeld & szymanik 2019, 2020, steinert-threlkeld 2020).12 our production paradigm also allows us to address the first key limitation. to do this, we enrich the input to our transformer. in addition to a mind and world representation, the input also tells the model what we want it to produce. put roughly, we either ask the network to produce the most informative true ascription, or ask for the most informative merely false ascription, or ask for the most informative ascription that is a presupposition failure.13 11this does not model the kind of comprehension one has if one grasps the truth conditions of an utterance, but lacks information that allows one to decide whether an ascription is true, or has such information but is somehow blocked from exploiting it. still, we take it to give us at least some access to how well one grasps the truth conditions of an utterance, notwithstanding concerns about the relation between competence and performance (compare dupre 2021). although our latest experiment tests production, similar points apply here too. 12johnson et al. (2021) investigated morphological universals for affixes using a production-based experiment with lstms as well as human subjects. outside the computational literature, maldonado & culbertson (2019) investigated semantic universals for person systems using a production-based experiment with human subjects. 13we thus test for presupposition failure without requiring our transformer to learn how to comprehend sentences used in the family of sentences diagnostic (attitude ascriptions embedded under negation, possibility, and question). proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 398 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ whether we ask for a truth, falsehood, or presupposition failure, we always ask for the most informative ascription. this ‘informativity demand’ highlights that the inputs to our transformer encode some pragmatic principles. since we effectively ask our transformer to maximize the informativity of asserted and presupposed content, our transformer learns how to satisfy grice (1989)’s maxim of quantity and heim (1991)’s maximize presupposition.14 this is key to getting our transformer to produce contrafactive and factive attitude ascriptions. because both types of ascription entail their non-factive counterpart, the model could maximize its chances of producing true ascriptions by producing non-factive attitude ascriptions only. our informativity demand forbids this behavior, since contrafactive and factive attitude ascriptions are more informative than their non-factive counterparts. given this key role of pragmatic principles in our new inputs, we label the inputs we use ‘semantic-pragmatic conditions.’ our new paradigm addresses both key limitations at once. this is more economical than addressing them separately. but it also raises a concern. we build in presupposition failure by sometimes asking our model to produce presupposition failures. but human language learners are hardly ever, if at all, asked to produce presupposition failures. so, our new paradigm is not as naturalistic as we would want. although we will weight our data to partially address this concern, we feel its force. to address it, we plan to do follow-up experiments that test for presupposition failure without asking for the production of presupposition failures. 3. current experiment. let’s turn to the details of our latest experiment.15 3.1. architecture. our transformer closely follows vaswani et al. (2017)’s models using the pytorch implementation. thus, our model consists of a transformer encoder and decoder. we previously used an encoder-only approach. but this is only appropriate for testing comprehension.16 to initialise the weights, we use the xavier uniform distribution. we 0-initialise biases and use sinusoidal position embeddings. but we use position-specific linear layers to constrain the output for each position in an output sequence to a proper subset of the words our model learns.17 in effect, our model learns an artificial language with a fixed word order and can only produce sentences with that word order. this means we can rule out the word order confound in experiment 1, whilst letting our model learn a language that is more naturalistic (due to its larger vocabulary) than in experiment 2. minimizing the role of syntax also allows our model to learn our artificial language more easily, rendering the training more economical.18 14our transformer also learns how to satisfy grice’s maxim of quality because it learns to produce a true ascription when we ask for one. of course, it also learns to produce falsehoods or presupposition failures when we ask for them. but this does not mean that it violates the maxim of quality because our demand that it produce something other than a true ascription suspends that maxim. 15the code for running our experiments as well as the output of our experiments can be found at https:// github.com/dstrohmaier/productive_contrafactives. 16the third and as yet unexplored option is a decoder-only approach, as popularised by the gpt-family of models (radford et al. 2018). 17one may worry that we are not testing production, as this constraint makes our task akin to a multiple classification task. however, even without that constraint, our task would be akin to a multiple classification task. the difference would merely be that, in classification, at a given position all words, and not just a proper subset of them, would have to be considered. more generally, there is no strict division between neural language modelling and classification. 18one may worry about minimizing the role of syntax. for cross-linguistically the meaning of attitude ascriptions proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 399 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3.2. input and output data. we train our transformer on a sequence-to-sequence task. it takes sequences of tokens and outputs sequences of tokens. the input sequences provide our semantic-pragmatic conditions: more specifically, what we call the ‘main value’, ‘sub value’, ‘mind-world relation’, and ‘attitude content’. the output sequences can be glossed as attitude ascriptions in an artificial language, consisting of an attitude verb and an embedded clause. the function from input to output can be written as:19 main value × sub value × mind-world relation × attitude content → attitude verb × embedded clause the first two inputs tell the model what we want it to produce. the main value is the value we want for the attitude ascription the model produces: true, false, or p-failure. the sub value is the value we want for the embedded clause the model uses in its attitude ascription: true, false, or unknown.20 the second two inputs give the model the information it needs to decide how to produce what we want it to produce. the mind-world relation has three possible values: mind and world match (=), are incompatible (!=), or the world state is unknown (?). finally, the attitude content exhaustively tells the model what the matrix subject believes. table 5 in the appendix lists all possible combinations of semantic-pragmatic conditions (in particular, mind-world relation, main value, and sub value) alongside their required attitude verbs.21 but let’s walk through three examples. say we have some attitude content, mind-world relation =, main value true, and sub value true. here (row 1), our model is required to produce a factive. or, say we have some attitude content, mind-world relation =!, main value true, and sub value false. now (row 11), our model is required to produce a contrafactive. or, say we have some attitude content with mind-world relation ?, main value true, and sub value unknown. in this case (row 21), our model is required to produce a non-factive. we let the sub value vary independently of main value and mind-world relation to permit the model to produce a wider range of attitude ascriptions. as mentioned earlier, a non-factive attitude ascription is less informative than its contrafactive and factive counterparts. so our model would often depends on syntactic properties. for instance, ozyildiz (2017) shows that whether turkish bil‘know’ has a truth presupposition depends on whether it embeds a nominalized clause or a tensed clause headed by diye. similar ‘factivity alternations’ are attested across many languages (bondarenko 2019, lee 2019, grano & park 2022). examples of dependence on syntactic properties can also be found in english (karttunen 1971). whether forget has a truth presupposition depends on whether it embeds a finite that-clause or a to-infinitive. a similar ‘factive-implicative’ alternation is also attested for remember. these examples are not problematic for our architecture because the artificial language our model learns is highly constrained. any language with factive attitude ascriptions has a type of embedded clause that gives us a truth presupposition. and one way to understand our artificial language is as containing just that type of embedded clause. this means that insofar as our transformers do not learn the factivity alternation, or the factive-implicative alternation, they do not learn the full meaning of predicates like biland forget. but we do not intend our transformers to do so. we focus on two dimensions along which contrafactives and factives are at least sometimes the same (belief entailment) and different (truth/falsity presupposition), and set aside potential alternations in which they participate. 19we use ‘function’ loosely here, as some inputs permit more than one kind of output. see below. 20although some tokens the model uses in the embedded clause can be glossed in english as presupposition triggers, e.g. ‘rory’, we set aside the possibility that the embedded clause is a presupposition failure. in effect, our transformer learns a language whose only presupposition triggers are contrafactives and factives. 21our data does not contain any of the rows marked as impossible. proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 400 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ not be permitted to produce a non-factive if we just asked for the main value false and gave it some mind-world relation. but by also asking for the sub value unknown, we can require the model to produce a false non-factive attitude ascription, even given our informativity demand. an input sequence generally requires exactly one output sequence. sometimes, though, the model has more than one option. table 5, row 27 gives an example: if we have some attitude content, mind-world relation ?, main value presupposition failure, and sub value unknown, both a contrafactive and a factive are permitted. we use slightly different languages for the input attitude content and the output embedded clause. table 1 lists the vocabulary used for the attitude content. this vocabulary gets us contents like ‘eat rory tomato basil soup lunch tomorrow’ or ‘buy ahab carrot oregano pie dinner yesterday’. the function from input attitude content to output embedded clause can then be written as: verb × agent × ingredient × spice × dish × meal × day → × agent × verb (with tense) × main ingredient + spice × preposition × dish × meal × day to make the task more naturalistic our mapping from input to output language is slightly indirect. the ‘+’ indicates that the ingredient and spice combine to a single output token. and the output verbs’ tense must be inferred from the input day indexical, e.g. the ‘tomorrow’ token. category lexical items verb eat, cook, order, buy subject rory, lorelai, lane, paris, timon, ahab ingredient tomato, pumpkin, mushroom, carrot, potato spice basil, oregano, pepper, chili, coconut dish soup, pie, rice, stew, curry meal lunch, dinner, breakfast day day-before-yesterday, yesterday, now, today, tomorrow, day-after-tomorrow table 1: attitude content vocabulary. content has one token of each category. an attitude content and embedded clause match iff applying the attitude content to the function from attitude content to embedded clause results in the embedded clause. for instance, ‘eat rory tomato basil soup lunch tomorrow’ matches ‘rory will-eat tomato-basil soup for lunch tomorrow’, but not ‘ahab bought carrot-oregano pie for dinner yesterday’. some semantic-pragmatic conditions require attitude content and embedded clause to match, some require them not to. to illustrate, the only way to produce a merely false non-factive attitude ascription is to use a nonmatching embedded clause (see table 5, rows 6, 15, and 24). this is because the input attitude content exhaustively describes what the matrix subject believes. so, any non-factive attitude ascription with a matching embedded clause is true, and any with a non-matching one false. 3.3. data generation. we generated input and output sequences automatically. because not all types of sequence would occur equally frequently if we generated all possible combinations in the same number (factives, in particular, would be over-represented amongst the output sequences), we subsampled to get a more balanced dataset. we also weighted the main semantic value, while ensuring that the sub value remained balanced. as noted earlier, human audiences rarely ask one to produce a sentence that is merely false proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 401 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ or a presupposition failure. to partially address this, whilst still allowing our model to learn how to produce such outputs, the number of instances where the main value is true (and we thereby ask for a true ascription) equals the sum of the number of instances where the main value is false and the number of instances where the main value is presupposition failure. the balanced data set contained 194,400 instances. for each attitude verb, 70,200 instances permit that verb. 16,200 instances permit both factives and contrafactives.22 113,400 instances require matching embedded clauses; 81,000 non-matching ones. we randomly split the balanced data into 90% training and 10% test data. 3.4. training. we split the training phase into two parts. the first was a hyperparameter search to identify suitable settings for hyperparameters, such as learning rate, size of training batches etc. the second was the training of the selected settings on all training data. we used the adamw algorithm (loshchilov & hutter 2018) with standard settings for the pytorch library (except for the learning rate, which is explored as a hyperparameter) and an adapted version of pytorch’s binary cross entropy (bce) loss. given semantic-pragmatic conditions that require a matching embedded clause, we set all output tokens (attitude verbs plus embedded clause constituents) required by those conditions to 1, the rest to 0. for conditions that permit any nonmatching clause, but not the matching one, we instead set the tokens that occur in the impermissible clause to 0.5. this reduces the likelihood that the model produces all these tokens together, while still allowing the model to produce individual ones. we then calculated the loss as the divergence from these values using the standard pytorch formula for bce. this resembles standard language modelling practice (e.g. the use of cross entropy loss in kaplan et al. 2020), but permits more than one output sequence per combination of semantic-pragmatic conditions. the hyperparameter search explored 41 different settings using a randomised search and 5fold cross-validation. our model learned how to produce our target expressions in 2 of these settings.23 table 2 details the search space and successful settings, which we call settings 1 and 2. name space setting 1 setting 2 dim. embedding {80, 100, 120, 140, 160, 180, 200} 180 200 dim. hidden {160, 180, 200, 220, 240, 260, 280, 300, 320, 340} 240 260 # attention heads {2, 5, 10, 20} 10 20 # encoder layers {5, 10, 15, 20, 25} 5 5 # decoder layers {5, 10, 15, 20, 25} 15 20 epochs {3, 5, 7} 3 5 batch size {120, 240, 360, 480, 600} 120 240 learning rate {1 ·10−3, 5 ·10−4, 1 ·10−4, 5 ·10−5, 1 ·10−5, 5 ·10−6, 1 ·10−6} 1 · 10−4 1 · 10−4 dropout {0.1, 0.2, 0.3} 0.1 0.3 table 2: hyperparameter space and two successful settings. 22one might worry that this gives a slight advantage to learning non-factives, as they never have to compete. we find, however, that our transformer learns non-factives most slowly. see below. 23the high failure rate in the search may raise concerns about the robustness of our results. however, neural networks regularly fail to learn throughout much of the hyperparameter space. our transformer is also relatively small, which leads us to expect that, unlike larger models, it only learns given relatively specific hyperparameters. proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 402 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3.5. evaluation. to track the learning process, we evaluated settings 1 and 2 on the complete test data every 20 training batches.24 this allowed us to compare how well our transformer had learned our target expressions at different stages of the training process. we also varied the original random seed for each setting four times to get 10 evaluations overall. by varying the random seed, we checked whether our results are robust to a random change in our network’s initial conditions. our evaluation metric was the correctness of the output sequence given the semantic-pragmatic conditions. if a condition permits contrafactive and factives, the output is correct with either verb. 3.6. results and discussion. figure 1 and table 3 show that we find no consistent difference between contrafactives and factives. while there are some differences, these reverse over the course of training. for instance, at setting 1 batch #200, the model performs better on factives (factives: 40.9%, contrafactives: 36.2%), but by batch #600 the order has reversed. figure 1: performance over the course of training. batch factive contrafactive non-factive 0 39.0 (±14.5) 26.5 (±17.4) 0.6 (±1.4) 200 40.9 (±5.9) 36.2 (±6.3) 12.8 (±8.4) 400 44.4 (±8.7) 41.8 (±9.0) 29.2 (±5.0) 600 65.5 (±3.1) 71.9 (±2.9) 56.1 (±4.9) 800 90.7 (±0.6) 94.2 (±1.8) 94.1 (±2.9) 1000 98.6 (±1.1) 99.5 (±0.3) 99.9 (±0.1) 1200 99.9 (±0.1) 99.9 (±0.1) 100.0 (±0.0) 1400 100.0 (±0.0) 100.0 (±0.0) 100.0 (±0.0) (a) setting 1. factive contrafactive non-factive 0 25.1 (±18.3) 29.8 (±21.9) 4.8 (±9.2) 400 48.5 (±5.1) 25.5 (±6.3) 0.5 (±0.6) 800 37.0 (±6.4) 41.0 (±8.8) 17.4 (±7.5) 1200 55.4 (±2.2) 52.7 (±6.6) 27.1 (±2.7) 1600 64.8 (±7.4) 65.1 (±6.7) 40.8 (±12.3) 2000 83.8 (±13.3) 83.9 (±13.8) 74.2 (±23.8) 2400 98.8 (±2.2) 98.9 (±2.1) 98.8 (±2.3) 2800 100.0 (±0.0) 100.0 (±0.0) 100.0 (±0.0) (b) setting 2. table 3: percentage of correct output sequences given permitted attitude verb. numbers in parentheses give standard deviation across random seeds. 24settings 1 and 2 differ in batch size (see table 2). nonetheless, comparison is relatively straightforward, as the batch size of setting 2 is exactly twice that of setting 1. proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 403 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (a) matching. (b) non-matching. figure 2: performance over the course of training split by whether model had to produce matching or non-matching embedded clauses. shaded area gives standard deviation between random seeds. non-factives. in early training, our model performs worst on non-factives. for example, at setting 2 batch #1,200, the model only treats correctly 17.4% of instances that require a nonfactive, but other attitude verbs perform above 50%. this parallels results from experiment 1. but the reason why we get this result is specific to our current experiment. some semantic-pragmatic conditions that require non-factives require non-matching embedded clauses. these are the conditions where non-factives perform worst. figure 2 shows that with matching embedded clauses non-factives perform no worse than contrafactives and factives. for setting 1 non-factives even perform better. to see why non-factives underperform with non-matching embedded clauses, consider nonmatching embedded clauses more generally. figure 2 shows that semantic-pragmatic conditions that require non-matching embedded clauses yield higher variability in early training and account for much of the variability in early training in figure 1. given the model’s freedom to choose which embedded clause to produce if a non-matching one is required, this may seem surprising. but, in fact, the dynamics are what we expect from a transformer. because only one embedded clause is matching, but 129,599 are non-matching, the model initially fares better with non-matching embedded clauses. but, as the model learns to produce matching clauses, it overadapts and produces them also in inappropriate cases. this leads to a correction in the other direction, giving rise to variability. in effect, the model is pushed every which way by the instances it learns from. the model’s overadaptation to matching clauses produces an even worse outcome for nonfactives than for contrafactives and factives. only one out of the four semantic-pragmatic conditions that require non-factives requires a matching embedded clause: table 5, row 21. but because of how we balanced the data (weighting main value true and balancing the verbs) this condition occurs much more frequently than conditions that require a non-factive with a non-matching embedded clause: 54,000 instances require a non-factive with a matching embedded clause, 16,200 a non-factive with a non-matching one.we do not get this imbalance with contrafactives and factives, for which 32,400 instances require matching embedded clauses and 37,800 require non-matching ones. this difference between non-factives on one side and contrafactives and factives on the other means that the model sees more conditions that require non-factives with matching embedded clauses than conditions that require contrafactives or factives with matching embedded clauses. proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 404 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ and this leads the model to overadapt to matching clauses more strongly for non-factives. hence, in our setup, non-factives underperform with non-matching embedded clauses. in effect, non-factives underperform because of the proportion of matching to non-matching embedded clauses that our model has to produce with non-factives. this proportion in turn results from our efforts to weight and balance semantic-pragmatic conditions and attitude verbs. we would, therefore, not be surprised if non-factives performed better given appropriate changes to how we weight and balance our inputs and outputs. selection preferences. some semantic-pragmatic conditions permit both contrafactives and factives. does our transformer prefer one verb over the other in these conditions? if so, this could indicate that the preferred verb is easier to learn. to answer this question, we look at selection preferences after the model has stabilised (batch >3,000). for every 20 batches after this threshold and every random seed, we consider the proportion of contrafactives and factives produced in semantic-pragmatic conditions that permit both. we focus on what happens after the model has stabilised because we expect that if one verb is easier to learn, the model would settle on producing that verb as a local optimum. and that would mean that it would consistently prefer that verb over the other once it stabilizes. table 4 and figure 3 show that the selection preference depends on the hyperparameters, not the attitude verb. our transformer does not prefer one verb over the other. setting verb count mean std min 25% 50% 75% max 1 contrafactive 345 0.76 0.13 0.28 0.68 0.77 0.85 0.99 factive 345 0.24 0.13 0.01 0.15 0.23 0.32 0.72 2 contrafactive 165 0.37 0.15 0.05 0.25 0.37 0.48 0.76 factive 165 0.63 0.15 0.24 0.52 0.63 0.75 0.95 table 4: statistics for selection (with batch >3,000). count (number of evaluations) differs between settings as batch sizes differ. numbers given are for the mean selected proportion, the standard deviation, and the quantiles for each setting and verb. figure 3: selection preference when both contrafactives and factives are permitted. 4. conclusion. natural languages appear to universally have factives, but lack contrafactives. we previously hypothesised that this asymmetry arises partly because the meaning of a contrafactive is harder to learn than that of a factive. our previous comprehension-oriented experiments supported this hypothesis, but the results from the production-oriented experiment reported here do not. proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 405 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ to fully interpret our results, more work is needed. given our results, the meaning of contrafactives may be harder to learn than that of factives in comprehension-oriented settings only. this would raise the question whether a difference in learnability in one setting is enough to explain any of the difference in frequency between contrafactives and factives. alternatively, it may be that our latest experiment has limitations that undermine the validity of its results. in fact, we highlighted one potential limitation earlier on: we built in presupposition failure by sometimes asking our model to produce presupposition failures. but human language learners are hardly ever, if at all, asked to produce presupposition failures. so, although we did ask for twice as many true ascriptions as presupposition failures, our latest experiment may not allow us to draw conclusions about human language learners.25 another potential limitation, compared to experiment 2 in particular, is that our current set-up does not capture how human infants seem to learn attitude verb meanings. on the pragmatic syntactic bootstrapping model (hacquard & lidz 2022), they infer the meaning of non-factive think, for instance, partly from the parallel between the use of non-factive attitude ascriptions like ayesha thinks that rebecca swims as indirect assertions and the primary use of their declarative complements as direct assertions. experiment 2 captured this developmental priority of complements over attitude verbs. we pre-trained our model to learn the meanings of embedded clauses; only once it learnt those meanings, did we train it on our attitude verbs. our current experiment, however, does not capture this developmental priority. this may be (part of) why our current results differ from those of experiment 2, and may mean that our current results do not give us as much insight into human language learners as we would want. in sum, given the results and (potential) limitations of our three experiments to date, more work remains to be done to assess whether a difference in learnability can contribute to an explanation of the difference in frequency between contrafactives and factives. 25one may also worry that our current experiment treats the difference between contrafactives and factives differently than our previous experiments. factive attitude ascriptions are true only if mind and world representations correspond, contrafactive attitude ascriptions only if they do not. our previous experiments treated these conditions asymmetrically. while there was only one world representation that corresponded to the mind representation, there were many world representations that failed to correspond. this asymmetry, the worry goes, may have made it harder to learn the meaning of contrafactives than that of factives. yet our current experiment treats the conditions symmetrically. both correspond to exactly one mind-world relation: = and !=. so, by making symmetrical what was asymmetrical, we may have removed a significant explanatory factor. and insofar as this factor is naturalistically motivated, our current experiment may run up against another limitation. in reply, note that we also levelled out an asymmetry in the other direction. factive attitude ascriptions are not true if mind and world representations do not correspond, whilst contrafactive attitude ascriptions are not true if mind and world representations do correspond. our previous experiments treated these conditions asymmetrically: only one world representation corresponded to the mind representation, but many world representations failed to correspond. by contrast, in our current experiment, both conditions correspond to exactly one mind-world relation: != and =. now, if, in our previous experiments, the first asymmetry made it harder to learn the meaning of a contrafactive than that of a factive, the second asymmetry may have made it harder to learn the meaning of a factive than that of a contrafactive. these asymmetries would then have balanced out, given that true and not true ascriptions were balanced in training and test data. in effect, our current experiment treats contrafactives and factives symmetrically from the start, whilst our previous experiments treat them asymmetrical twice and in opposite, and therefore balancing, directions. ultimately, we expect no significant difference between these treatments. proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 406 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ references anvari, amir, mora maldonado & andrés soria ruiz. 2019. the puzzle of reflexive belief construction in spanish. proceedings of sinn und bedeutung 23(1). 57–74. https://doi.org/10.18148/sub/2019.v23i1.503. bochnak, m. ryan & emily a. hanink. 2022. clausal embedding in washo: complementation vs. modification. natural language & linguistic theory 40(4). 979–1022. https://doi.org/10.1007/s11049-021-09532-z. bondarenko, tatiana. 2019. from think to remember: how cps and nps combine with attitudes in buryat. semantics and linguistic theory 29. 509–528. https://doi.org/10.3765/salt.v29i0.4605. bossi, madeline. 2022. unifying negative bias and reminding functions: the case of kipsigis par. in özge bakay, breanna pratley, evan neu & peyton deal (eds.), proceedings of the fifty-second annual meeting of the north east linguistic society, 95–104. amherst. caucheteux, charlotte & jean-rémi king. 2022. brains and algorithms partially converge in natural language processing. communications biology 5(1). https://doi.org/10.1038/s42003022-03036-1. dupre, gabe. 2021. (what) can deep learning contribute to theoretical linguistics? minds and machines 31(4). 617–635. https://doi.org/10.1007/s11023-021-09571-w. von fintel, kai. 2004. would you believe it? the king of france is back! (presuppositions and truth-value intuitions). in marga reimer & anne bezuidenhout (eds.), descriptions and beyond, 269–296. oxford: clarendon press. goddard, cliff. 2010. universals and variation in the lexicon of mental state concepts. in words and the mind: how words capture human experience, oxford: oxford university press. grano, thomas & jisu park. 2022. to the best of our knowledge: factivity alternation in korean. in özge bakay, breanna pratley, evan neu & peyton deal (eds.), proceedings of the fifty-second annual meeting of the north east linguistic society, 307–316. amherst. https://www.dropbox.com/s/r76p8ir3mhmd1bd/granoparknels52v3.pdf?dl=0. grice, paul. 1989. studies in the way of words. cambridge, mass.: harvard university press. hacquard, valentine & jeffrey lidz. 2022. on the acquisition of attitude verbs. annual review of linguistics 8(1). 193–212. https://doi.org/10.1146/annurev-linguistics-032521-053009. hannon, michael. 2015. the universal core of knowledge. synthese 192(3). 769–786. https://doi.org/10.1007/s11229-014-0587-y. heim, irene. 1991. artikel und definitheit. in semantik, 487–535. berlin: de gruyter. publisher: de gruyter. hoeksema, jack. 2021. verbs of deception, point of view and polarity. proceedings of the international conference on head-driven phrase structure grammar 26–46. https://doi.org/10.21248/hpsg.2021.2. holton, richard. 2017. i—facts, factives, and contrafactives. aristotelian society supplementary volume 91(1). 245–266. https://doi.org/10.1093/arisup/akx003. hsiao, pei-yi katherine. 2017. on counterfactual attitudes: a case study of taiwanese southern min. lingua sinica 3(1). 4. https://doi.org/10.1186/s40655-016-0019-7. hyman, john. 2017. ii—knowledge and belief. aristotelian society supplementary volume 91(1). proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 407 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 267–288. https://doi.org/10.1093/arisup/akx005. johnson, tamar, kexin gao, kenny smith, hugh rabagliati & jennifer culbertson. 2021. investigating the effects of i-complexity and e-complexity on the learnability of morphological systems. journal of language modelling 9(1). 97–150. https://doi.org/10.15398/jlm.v9i1.259. kallini, julie, isabel papadimitriou, richard futrell, kyle mahowald & christopher potts. 2024. mission: impossible language models. in lun-wei ku, andre martins & vivek srikumar (eds.), proceedings of the 62nd annual meeting of the association for computational linguistics, 14691–14714. bangkok, thailand: association for computational linguistics. https://aclanthology.org/2024.acl-long.787. kaplan, jared, sam mccandlish, tom henighan, tom b. brown, benjamin chess, rewon child, scott gray, alec radford, jeffrey wu & dario amodei. 2020. scaling laws for neural language models https://doi.org/10.48550/arxiv.2001.08361. karttunen, lauri. 1971. implicative verbs. language 47(2). 340–358. https://doi.org/10.2307/412084. kierstead, gregory weiss. 2015. projectivity and the tagalog reportative evidential: the ohio state university ma thesis. https://etd.ohiolink.edu/apexprod/rwsolink/r/1501/10?p10etdsubid = 106688clear = 10. krifka, manfred. 2016. realis and non-realis modalities in daakie (ambrym, vanuatu). semantics and linguistic theory 566–583. https://doi.org/10.3765/salt.v26i0.3865. lee, chungmin. 2019. factivity alternation of attitude ‘know’ in korean, mongolian, uyghur, manchu, azeri, etc. and content clausal nominals. journal of cognitive science 20(4). 451– 503. https://doi.org/10.17791/jcs.2019.20.4.451. loshchilov, ilya & frank hutter. 2018. decoupled weight decay regularization, https://openreview.net/forum?id=bkg6ricqy7. maldonado, mora & jennifer culbertson. 2019. learnability as a window into universal constraints on person systems. proceedings of the amsterdam colloquium 22. maldonado, mora, jennifer culbertson & wataru uegaki. 2022. learnability and constraints on the semantics of clause-embedding predicates. https://doi.org/10.31234/osf.io/zst5y. mcgregor, william b. 2024. on the expression of mistaken beliefs in australian languages. linguistic typology 28(1). 101–145. https://doi.org/10.1515/lingty-2022-0023. merkx, danny & stefan l. frank. 2021. human sentence processing: recurrence or attention? in proceedings of the workshop on cognitive modeling and computational linguistics, 12–22. association for computational linguistics. https://doi.org/10.18653/v1/2021.cmcl-1.2. ozyildiz, deniz. 2017. attitude reports with and without true belief. semantics and linguistic theory 27(0). 397–417. https://doi.org/10.3765/salt.v27i0.4189. paape, dario. 2023. when transformer models are more compositional than humans: the case of the depth charge illusion. experiments in linguistic meaning 2. 202–218. https://doi.org/10.3765/elm.2.5370. pham, thang, trung bui, long mai & anh nguyen. 2021. out of order: how important is the sequential order of words in a sentence in natural language understanding tasks? in findings of the association for computational linguistics: acl-ijcnlp 2021, 1145–1160. association for computational linguistics. https://doi.org/10.18653/v1/2021.findings-acl.98. proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 408 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ phillips, jonathan & aaron norby. 2021. factive theory of mind. mind & language 36(1). 3–26. https://doi.org/10.1111/mila.12267. radford, alec, karthik narasimhan, tim salimans & ilya sutskever. 2018. improving language understanding by generative pre-training . roberts, tom & deniz özyildiz. 2023. bad attitudes. rosenberg, marc stephen. 1975. counterfactives: a pragmatic analysis of presupposition: university of illinois at urbana-champaign dissertation. ross, alexis & ellie pavlick. 2019. how well do nli models capture verb veridicality? in kentaro inui, jing jiang, vincent ng & xiaojun wan (eds.), proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlp-ijcnlp), 2230–2240. hong kong, china: association for computational linguistics. https://doi.org/10.18653/v1/d19-1228. sander, thorsten. 2020. fregean side-thoughts. australasian journal of philosophy 0(0). https://doi.org/10.1080/00048402.2020.1795216. schrimpf, martin, idan asher blank, greta tuckute, carina kauf, eghbal a. hosseini, nancy kanwisher, joshua b. tenenbaum & evelina fedorenko. 2021. the neural architecture of language: integrative modeling converges on predictive processing. proceedings of the national academy of sciences 118(45). https://doi.org/10.1073/pnas.2105646118. steinert-threlkeld, shane. 2020. an explanation of the veridical uniformity universal. journal of semantics 37(1). 129–144. https://doi.org/10.1093/jos/ffz019. steinert-threlkeld, shane & jakub szymanik. 2019. learnability and semantic universals. semantics and pragmatics 12(0). https://doi.org/10.3765/sp.12.4. steinert-threlkeld, shane & jakub szymanik. 2020. ease of learning explains semantic universals. cognition 195. https://doi.org/10.1016/j.cognition.2019.104076. strohmaier, david & simon wimmer. 2022. contrafactives and learnability. in marco degano, tom roberts, giorgio sbardolini & marieke schouwstra (eds.), proceedings of the 23rd amsterdam colloquium, 298–305. amsterdam. https://www.dropbox.com/s/umjf5rn8mjq3rbx/proceedings2022.pdf?dl=0. strohmaier, david & simon wimmer. 2023. contrafactives and learnability: an experiment with propositional constants. in daisuke bekki, koji mineshima & eric mccready (eds.), logic and engineering of natural language semantics lecture notes in computer science, 67–82. cham: springer nature switzerland. https://doi.org/10.1007/978-3-031-43977-3 5. vaswani, ashish, noam shazeer, niki parmar, jakob uszkoreit, llion jones, aidan n gomez, łukasz kaiser & illia polosukhin. 2017. attention is all you need. 31st conference on neural information processing systems 1–11. wimmer, simon bastian. 2019. reflections on knowledge and belief : university of warwick phd. http://webcat.warwick.ac.uk/record=b3494900 s1. ziembicki, daniel, karolina seweryn & anna wróblewska. 2023. polish natural language inference and factivity: an expert-based dataset and benchmarks. natural language engineering 1–32. https://doi.org/10.1017/s1351324923000220. proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 409 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ appendix mind-world main value sub value attitude verb embedded clause 1 = true true factive matching 2 = true false impossible — 3 = true unknown impossible — 4 = false true factive non-matching 5 = false false contrafactive non-matching 6 = false unknown non-factive non-matching 7 = p-failure true contrafactive matching 8 = p-failure false factive non-matching 9 = p-failure unknown factive or contrafactive non-matching 10 != true true impossible — 11 != true false contrafactive matching 12 != true unknown impossible — 13 != false true factive non-matching 14 != false false contrafactive non-matching 15 != false unknown non-factive non-matching 16 != p-failure true contrafactive non-matching 17 != p-failure false factive matching 18 != p-failure unknown factive or contrafactive non-matching 19 ? true true impossible — 20 ? true false impossible — 21 ? true unknown non-factive matching 22 ? false true factive non-matching 23 ? false false contrafactive non-matching 24 ? false unknown non-factive non-matching 25 ? p-failure true contrafactive non-matching 26 ? p-failure false factive non-matching 27 ? p-failure unknown factive or contrafactive matching table 5: possible combinations of semantic-pragmatic conditions and output tokens. proceedings of elm 3: 395-410, 2025 david strohmaier and simon wimmer: contrafactives, learnability, and production. 410 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ you must worry! the interpretation of mustn’t varies with context and verbal complement adina camelia bleotu, anton benz & roxana pǎtrunjel* abstract. we investigate experimentally whether american english (ae) adult speakers are influenced in their interpretation of mustn’t by pragmatic context (contexts favoring lack of necessity/necessity not to readings) and/or the semantic properties of the verbal complements of the modal (verbs denoting events in the physical realm vs. verbs expressing undesirable mental activities). in an experiment combining a forced choice task and a gradient acceptability task, participants saw sentences containing mustn’t and physical events/negative mental activities in lack of necessity/necessity not to contexts (e.g., you mustn’t worry. the woman will give you money) they had to choose the most suitable interpretation of mustn’t (‘it is necessary not to’/ ‘it is not necessary’ interpretations). they then had to rate the acceptability of the sentences containing mustn’t in context on a likert scale from 1 to 7. we find that participants split into two groups: an interdiction group, which always treated mustn’t as expressing interdiction, and a variation group, which tended to interpret mustn’t as lack of necessity when the context favored such a reading and when the verbal complement the modal combined with was a negative mental activity. we argue that the lack of necessity reading of mustn’t is obtained via pragmatic weakening from its primary interdiction reading, and that this process is sensitive to context, as well as to the cognitive difficulty of imposing or forbidding mental (but not physical) activities to others. keywords. american english; deontic modality; deontic necessity; negation; interdiction; context-sensitivity; mental activities 1. introduction. in the current paper, we investigate experimentally how mustn’t is interpreted in ae, probing into whether such interpretation may vary with pragmatic context and the semantic properties of the verb taken as a complement by the modal. while mustn’t is generally assumed to express interdiction (as in you mustn’t smoke), context and verbal semantics may lead to a weaker interpretation akin to needn’t. regarding pragmatic context, we are interested in whether ae speakers interpret mustn’t differently depending on whether the pragmatic context of the utterance favors a lack of necessity reading or a necessity not to reading, in a similar fashion to how speakers vary their interpretation of scalar items depending on context (ronai & xiang 2020). for instance, (1a) exemplifies a lack of necessity context: it is not necessary to eat the bread today, since it will be good to eat tomorrow as well. (1b), on the other hand, exemplifies a * this research was supported by the dfg project sigames2: experimental game theory and scalar implicatures: investigating variation in context and scale type led by dr. anton benz (grant nr. be 4348/4-2). adina camelia bleotu was supported by a visiting scholar fellowship in semantics and pragmatics at leibniz-zentrum allgemeine sprachwissenschaft (zas) during 1 october 2021-1 april 2022. benz was partially supported by the project erc 787929 spagad: speech acts in grammar and discourse. we are grateful to the audiences at the human sentence processing conference (24-26 march 2022, usantacruz) and experiments in linguistic meaning 2 (18-20 may 2022, upenn) for their useful comments and suggestions. authors: adina camelia bleotu, zas berlin, university of bucharest (cameliableotu@gmail.com), anton benz, zas berlin (benz@leibniz-zas.de), & roxana pǎtrunjel, university of bucharest (roxanamihaela2097@gmail.com). proceedings of elm 2: 24-35, 2023 c©2023 adina camelia bleotu, anton benz and roxana pǎtrunjel published by the lsa with permission of the author(s) under a cc by license. 24 https://doi.org/10.3765/elm https://www.elm-conference.net/ necessity not to context: it would be ideal not to eat the bread since there should be enough bread for the visitors when they come over. (1) a. tom mustn’t eat the bread. it won’t go stale by tomorrow. (lack of necessity context) b. tom mustn’t eat the bread. they have visitors coming over. (necessity not to context) regarding the semantic properties of the modal complement, we are interested in whether ae speakers interpret mustn’t differently depending on whether the lexical verb expresses a nonbeneficial/negative mental activity or a physical event. speakers may be more likely to interpret mustn’t as needn’t when the verbal complement of the modal expresses a negative mental activity, which may be more difficult to control than an event in the physical realm: (2) a. he mustn’t panic. the teacher will give the class an easy test. (lack of necessity context) b. he mustn’t panic. the bears will attack him otherwise. (necessity not to context) 2. background. 2.1. background on modality and negation in english. english modals display an irregular behaviour in interaction with negation. while negation has a fixed position in english (i.e., always after a modal), its semantics is variable (i.e., the negation may scope below/above modality). this is true for different modals (3 a, b), as well as different flavors of the same modal (deontic & epistemic-3 b, c) (palmer 1995). (3) a. the boy must not/mustn’t go to the party. (obligation>not) b. the girl may not in the park this evening. (not>permission) c. the girl may not be doing her homework. (possible>not) various attempts at generalizations have been put forth, either in terms of the possibility/necessity distinction (cormack & smith 2002), or in terms of the deontic/epistemic modality distinction (coates 1983, picallo 1990), but, as pointed out by these authors themselves, there are always exceptions to these generalizations. moreover, it is unclear what n’t and not represent from a syntactic point of view (sentence or adverbial negation). in the ideal situation, sentence negation translates as external negation (neg>modality), and adverbial negation translates as internal negation (modality>neg) — see palmer (1995), for instance. however, we find cases where what looks like sentence negation (n’t) expresses internal negation (see (3a)). 2.2. background on the meaning of must in interaction with negation. deontic must not and mustn’t are generally argued to express interdiction (coates 1983, palmer 1995, papafragou 2000, huddleston & pullum 2002, cormack & smith 2002 a.o). (4) you must not/mustn’t smoke. however, in special polarity-sensitive contexts like contrastive negation (israel 1996, homer 2011, iatridou & zeijlstra 2013, zeijlstra 2017, a.o.), deontic necessity scopes below negation: (5) no student must read 5 articles on the topic but one student is encouraged to do so. meaning “it is not necessary that all students read 5 articles on the topic but one student is encouraged to do so.” in polarity-neutral contexts, however, there seems to be general consensus that deontic must not and mustn’t lead to an interpretation where deontic necessity scopes above negation. in the case of must not, this interpretation could be explained by arguing that not is actually adverbial negation, negating the verbal complement rather than sentence negation. in the case of mustn’t, proceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 25 https://doi.org/10.3765/elm https://www.elm-conference.net/ however, such an explanation cannot hold, given that n’t is a clear marker of sentence negation. consequently, we are in a situation where there seems to be no transparent mapping between syntax and semantics. from a syntactic point of view, n’t negates the modal, while, from a semantic point of view, it negates the verbal complement of the modal. our question is whether mustn’t always expresses interdiction or whether it can also express lack of necessity in combination with certain verbs and in certain contexts. with this aim in mind, we conducted a search in the contemporary corpus of american english (coca, davies 2008) for examples from ae containing mustn’t, which are odd under an interdiction reading. we draw attention to cases such as (6), which are not (obviously) polarity-dependent, and which seem to involve lack of necessity contexts and negative mental activity verbs. (6) a. you mustn’t worry about money. i’ll cover all your expenses. b. don’t worry. i’ll always take care of you. you mustn’t worry about being old. c. every man is to find his own strength. when you meet hardships you mustn’t panic, no matter how big the challenge is. interestingly, mustn’t worry is less frequent than needn’t worry, which always expresses a lack of necessity meaning (mustn’t worry occurs in 38 examples in coca, whereas needn’t worry occurs in 248 examples). in the examples where it does occur with mental verbs such as worry (6), mustn’t seems to be interpreted more like needn’t, such that deontic necessity scopes below negation. since n’t is typically a marker of sentence negation, mustn’t may be more prone to a lack of necessity interpretation than must not, while not is ambiguous between sentence or adverbial negation depending upon stress. this idea seems to be supported by the fact that, must not worry is quite infrequent, being found only in 5 examples (see 7). (7) we must not worry about them. god is responsible for dealing with them. cases where mustn’t has a similar interpretation to needn’t are problematic for the semantics of mustn’t, and it is not clear whether the different scope relations obtaining between deontic necessity and negation in such cases are a consequence of the lack of necessity context, the semantics of the verbal complement or of both. 3. current experiment. to get a clearer picture of the effect of pragmatic context and the semantics of the verbal complement upon the interpretation of mustn’t, we decided to conduct an experiment on ae adult speakers where we explicitly manipulate these factors. 3.1. predictions. as far as context is concerned, if the scope of negation and deontic necessity in mustn’t is fixed, then we expect the lack of necessity interpretation in both types of contexts. if, however, the scope of negation and deontic necessity in the case of mustn’t is sensitive to context, lack of necessity contexts should lead to a lack of necessity reading of mustn’t, while necessity not to contexts should lead to a necessity not to reading. as far as verbal complement semantics is concerned, we predict that, from a cognitive perspective, it should be harder to forbid someone not to experience a certain negative mental state (like worry or panic) than to forbid them to do a certain action (like eating bread), because people have less control over their emotions than over their actions. thus, we expect participants to interpret mustn’t followed by negative mental activities as a suggestion rather than an interdiction. however, the opposite effect might be expected if one considers the contrast relation between mustn’t and needn’t, more specifically, the high frequency of needn’t in combination with negative mental states compared to mustn’t (needn’t worry occurs in 248 examples in coca, while mustn’t worry only in 38). proceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 26 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.2. participants using mechanical turk, we collected and analyzed answers from 34 ae adult speakers (age range: 18-63, mean age: 36;8). participants received monetary compensation for completing the experiment. 3.3. methodology the experiment was preceded by a warm-up, where participants read 4 acceptable or unacceptable sentences that contained either negative imperatives or the negative of have to. participants had to interpret such sentences in terms of paraphrases of the type ‘it is necessary not to’ or ‘it is not necessary’. they were also asked to judge how acceptable these sentences were to them on a likert scale from 1 to 7. importantly, half of the sentences represented an inadequate use of the imperative or have to in order to increase participants’ attentiveness. table 1 exemplifies one such warm-up sentence containing a negative imperative. table 1: sample warm-up items our experiment combined a forced choice task and a gradient acceptability task for 8 critical sentences containing mustn’t and 16 control sentences containing needn’t and shouldn’t. just as in the warm-up, participants had to first choose the most adequate interpretation of a sentence (either ‘it is necessary not to’ or ‘it is not necessary’), and then judge the acceptability of that sentence on a likert scale from 1 to 7. the critical sentences used different verb types (mental/physical) in different pragmatic contexts (lack of necessity/necessity not to). participants saw verbs occurring in only one context. for each verb type, 4 verbs were tested: negative mental verbs such as worry, panic, be sad, be upset and physical verbs such as eat, drink, do, speak. half of the sentences had 2nd person pronouns, and half 3rd person pronouns. in the critical conditions, participants read a sentence containing mustn’t and had to choose the most suitable interpretation (either ‘it is necessary not to’ or ‘it is not necessary’ interpretations — see table 2 and table 3 for examples involving both mental and physical activities). they then had to rate the acceptability of the sentence in context on a likert scale from 1 to 7. table 2: sample critical items for combining mustn’t with a mental activity verb forced choice task in don’t be tall! there are enough tall people in the room, the sentence don’t be tall! means: a. it is necessary that you are not tall. b. it is not necessary that you are tall. acceptability judgment task (on a likert scale from 1 to 7) how acceptable do you think the sentence don’t be tall! is in the context don’t be tall! there are enough tall people in the room? (fully unacceptable) 1 2 3 4 5 6 7 (fully acceptable) forced choice task lack of necessity context: in you mustn’t worry. the woman will give you money, the sentence you mustn’t worry means: necessity not to context: in you mustn’t worry. you will get sick otherwise, the sentence you mustn’t worry means: a. it is necessary that you do not worry. b. it is not necessary that you worry. acceptability judgment task (on a likert scale from 1 to 7) lack of necessity context: how acceptable do you think the sentence you mustn’t worry is in the context you mustn’t worry. the woman will give you money? necessity not to context: how acceptable do you think the sentence you mustn’t worry is in the context you mustn’t worry. you will get sick otherwise? (fully unacceptable) 1 2 3 4 5 6 7 (fully acceptable) proceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 27 https://doi.org/10.3765/elm https://www.elm-conference.net/ table 3: sample experimental items for combining mustn’t with a physical activity verb just as the critical sentences, the control sentences also used both types of verbs (mental/physical activities) in either acceptable or unacceptable contexts (see tables 4 and 5). just as the warm-up sentences, the controls represented a useful measure of participants’ attention. participants were expected to generally interpret sentences with needn’t as expressing a lack of necessity meaning, while interpreting sentences with shouldn’t as expressing a necessity not to meaning. moreover, participants were expected to rate odd sentences as less acceptable. table 4: sample control item with needn’t in an adequate lack of necessity context table 5: sample control item with needn’t in an inadequate necessity not to context 3.4. results. we first take a look at participants’ answers for control sentences. participants generally interpret controls with needn’t as expressing lack of necessity meanings (at a rate of 74.3% for pragmatically adequate sentences and 70.9% for sentences where needn’t is pragmatically inadequate). however, in terms of acceptability, they give pragmatically adequate sentences with needn’t a 6.05 rating, while they give pragmatically inadequate sentences with needn’t a lower rating (3.76). in terms of response times, following ronai & xiang (2020), we removed the trials with extremely long and short responses, eliminating the top and bottom 2.5% of the data. looking at the response times in the forced choice task, we find participants take longer to interpret pragmatically inadequate sentences with needn’t (8659ms) compared to forced choice task lack of necessity context: in you mustn't drink alcohol. you are already in good spirits, the sentence you mustn’t drink alcohol means: necessity not to context: in you mustn't drink alcohol. it will make you feel sick, the sentence you mustn’t drink alcohol means: a. it is necessary that you do not drink alcohol. b. it is not necessary that you drink alcohol. acceptability judgment task (on a likert scale from 1 to 7) lack of necessity context: how acceptable do you think the sentence you mustn’t drink alcohol is in the context you mustn't drink alcohol. you are already in good spirits? necessity not to context: how acceptable do you think the sentence you mustn’t drink alcohol is in the context you mustn’t drink alcohol. it will make you feel sick? (fully unacceptable) 1 2 3 4 5 6 7 (fully acceptable) forced choice task in tom needn’t be offended. the woman didn’t want to insult him at all, the sentence tom needn’t be offended means: a. it is necessary that tom is not offended. b. it is not necessary that tom is offended. acceptability judgment task (on a likert scale from 1 to 7) how acceptable do you think the sentence tom needn’t be offended is in the context tom needn’t be offended. the woman didn’t want to insult him at all? (fully unacceptable) 1 2 3 4 5 6 7 (fully acceptable) forced choice task in you needn’t sweep the floor. it is very dirty, the sentence you needn’t sweep the floor means: a. it is necessary that you sweep the floor. b. it is not necessary that you sweep the floor. acceptability judgment task (on a likert scale from 1 to 7) how acceptable do you think the sentence you needn’t sweep the floor is in the context you needn’t sweep the floor. it is very dirty? (fully unacceptable) 1 2 3 4 5 6 7 (fully acceptable) proceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 28 https://doi.org/10.3765/elm https://www.elm-conference.net/ pragmatically adequate sentences with needn’t (6496ms). looking at the response times in the acceptability task, we find that participants also take longer to judge the acceptability of pragmatically inadequate sentences with needn’t (2833ms) compared to pragmatically adequate sentences with needn’t (2314ms). as far as control sentences with shouldn’t are concerned, we find that participants mostly interpret them as expressing necessity not to meanings in pragmatically adequate contexts (at a rate of 60.81%), as well as in pragmatically inadequate contexts (at a rate of 65.54%). in terms of acceptability, they rate pragmatically inadequate sentences with shouldn’t as 3.74, a lower rating than for pragmatically adequate sentences (5.88). in terms of response times, they take slightly longer to provide interpretations for pragmatically inadequate sentences with shouldn’t (9772ms) than for pragmatically adequate ones (9217ms). moreover, they take longer to judge the acceptability of pragmatically inadequate sentences with shouldn’t (3189ms) compared to pragmatically adequate ones (2441ms). overall, the answers provided by participants in the control sentences suggest that participants were attentive when interpreting and evaluating the sentences they read. for this reason, we decided not to remove any participants from the data analysis. we next consider the answers provided by participants for the critical trials containing mustn’t. a look at the overall answers reveals a higher rate of necessity not to readings of mustn’t in necessity not to contexts compared to lack of necessity contexts (figure 1). moreover, in lack of necessity contexts, participants seem to interpret mustn’t as expressing a lack of necessity meaning more with mental verbs than with physical verbs. figure 1: necessity not to readings overall interestingly, participants found the use of mustn’t in context quite natural, since mustn’t was rated as very acceptable in both necessity not to (6.11) and lack of necessity contexts (5.44), with mustn’t being rated slightly higher in necessity not to contexts. this suggests that both necessity not to and lack of necessity readings are available to ae adult speakers. we also look at response times, keeping in mind, however, that they should be taken with caution. since we did not control for sentence length, contexts with physical verbs were sometimes longer, involving more words and transitive verbs, which may have affected the results in the forced choice task and the acceptability judgment task. in the forced choice task, participants took longer to interpret answers in lack of necessity contexts (8194ms) compared to necessity not to contexts (6232ms). in necessity not to contexts, participants took slightly longer to interpret mustn’t when combined with physical verbs (6650ms) compared to mental verbs (5813ms). in lack of necessity contexts, participants took longer to interpret mustn’t when combined with physical verbs (9342ms) than with mental verbs (7045ms). in the rating task, participants took longer to rate answers in lack of necessity contexts (3445ms) compared to the necessity not to contexts (3015ms). in necessity not to contexts, parproceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 29 https://doi.org/10.3765/elm https://www.elm-conference.net/ ticipants took slightly longer to rate mustn’t when combined with mental verbs (3093ms) compared to physical verbs (2938ms). in lack of necessity contexts, participants took longer to rate mustn’t when combined with physical verbs (4285ms) than with mental verbs (2606ms). in order to see whether the variation in the forced choice answers characterizes every participant or only some, we also looked at the answers provided by individual participants. we notice that all participants seem to converge in their interpretation of mustn’t as expressing a necessity not to meaning in necessity not to contexts, but that there seems to be interparticipant variation in the interpretation of mustn’t in lack of necessity contexts. 15 participants, who we shall refer to as the interdiction group, interpret mustn’t as expressing lack of necessity in lack of necessity contexts more than 75% of the time (see figure 2). 19 participants, who we shall refer to as the variation group, were more context-sensitive, providing more lack of necessity readings than necessity not to readings in lack of necessity contexts, especially with mental verbs (see figure 3). interestingly, the interdiction group judged sentences in necessity not to contexts as 5.95 on a likert scale from 1 to 7, while rating sentences in lack of necessity contexts as 5.01. in contrast, the variation group judged sentences in necessity not to contexts as 6.23 on a likert scale from 1 to 7, while rating sentences in lack of necessity contexts as 5.75. the variation group was thus more at ease with the use of mustn’t in both lack of necessity and necessity not to contexts, while the interdiction group judged the use of mustn’t in lack of necessity contexts to be somewhat less acceptable. fig. 2: necessity not rates (interdiction group) fig. 3: necessity not rates (variation group) in terms of the response times in the forced choice task, participants from the interdiction group took longer to interpret mental verbs compared to physical verbs in both necessity not to contexts (7508ms>6684) and lack of necessity contexts (7101ms>5937ms). participants from the variation group took longer to interpret lack of necessity contexts (5566ms) compared to necessity not to contexts (9411ms). in necessity not to contexts, the variation group took slightly longer to interpret mustn’t and physical verbs (5990ms) compared to mental verbs (5142ms). in lack of necessity contexts, the variation group took longer to interpret mustn’t and physical verbs (1109ms) compared to mental verbs (7813ms). in terms of likert rating times, the interdiction group was generally slower in rating mustn’t in combination with physical verbs than with mental verbs in both necessity not to contexts (4195ms>3063ms) and lack of necessity contexts (5146ms>2776ms). in contrast, the variation group was faster in rating mustn’t with mental verbs (2489ms) than with physical verbs (3644ms). however, in necessity not to contexts, the variation group was faster in rating mustn’t with physical verbs (1970ms) compared to mental verbs (3116ms). it is not clear whether the response times differences presented above bear statistical significance, but, at a first glance, both the response times in the forced choice task and the likert proceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 30 https://doi.org/10.3765/elm https://www.elm-conference.net/ rating times seem to suggest that the variation group is somewhat more sensitive to context and the type of verb compared to the interdiction group. in terms of statistical analysis, we first conducted an overall analysis of the data. using r 4.0.5 (2021), we fitted a generalized mixed effects model with answer in the forced choice interpretation task as a dependent variable (dv), coded as 1 if sensitive to context and 0 in case not, and verb type (mental/physical), context (necessity not to/lack of necessity) and their interaction as fixed effects and random slopes per item and participant. we found a significant effect for context (ß = 3.098, se = 0.851, z = 3.640, p <.01) and the interaction between verb type and context (ß = −1.926, se = 0.943, z = −2.040, p <.05). participants gave fewer necessity not to readings in lack of necessity contexts and with mental verbs. interestingly, in necessity not to contexts, there is no difference in answers depending on the verb type. to see whether acceptability ratings vary with context and verb type, we fitted a linear mixed effects model with likert rating as a dv and verb type (mental/physical), context (necessity not to/lack of necessity) and their interaction as fixed effects and random slopes per item and participant. we found a significant effect for the interaction between verb type and context (ß = −1.557, se = 0.348, df = 5.009, t = −4.463, p < .01). while there was no significant difference in ratings for mustn’t in combination with physical or mental verbs in necessity not to contexts (mustn’t + physical verb was rated 6.19, mustn’t + mental verb was rated 6.04), there was a significant difference in these ratings in lack of necessity contexts: participants rated mustn’t combined with physical verbs as 4.74, much lower than mustn’t combined with mental verbs (6.13) we also looked at response times for lack of necessity interpretations in the forced choice task: we ran a mixed effects regression with the logarithm of response times (log (rt)) as a dv, verb type (mental/physical), context (necessity not to/lack of necessity) and their interaction as fixed effects, and item and participant as random effects (as the random slopes models did not converge). the results reveal no significant effect. however, a parallel model run on the likert rating times for lack of necessity readings reveals a significant interaction between context and verb type (ß = 0.525, se = 0.217, df = 171.238, t = 2.413, p < .05). in necessity not to contexts, participants give faster ratings for sentences with physical verbs, but in lack of necessity contexts, they are faster with mental verbs. we also analyzed the data from the interdiction group and the variation group separately. we performed a logistic regression, fitting the interdiction group answers in the forced choice task into a generalized linear mixed effects model with verb type (mental/physical), context (necessity not to/lack of necessity) and their interaction as fixed effects and random slopes per item and participant. we found a significant effect of context (ß = −7.811, se = 3.183, z = −2.454, p < .05), but no other significant effect. to see whether context and/or verb type had an effect on acceptability ratings, we fitted the data into a linear mixed effects model with likert rating as a dv and verb type (mental/physical), context (necessity not to/lack of necessity) and their interaction as fixed effects and random slopes per item and participant. we found a significant effect for the interaction between verb type and context (ß = −1.49, se = 0.464, df = 15.563, t = −3.214, p < .01). while we found no significant difference in ratings for mustn’t in combination with physical or mental verbs in necessity not to contexts (mustn’t + physical verb was rated as 6, mustn’t + mental verb was rated as 5.9), there was a significant difference in these ratings in lack of necessity contexts: participants rated mustn’t combined with physical verbs as 4.23, much lower than mustn’t combined with mental verbs (5.8). in terms of response times for interpretation in the forced choice task, we separately ran two linear mixed effects reproceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 31 https://doi.org/10.3765/elm https://www.elm-conference.net/ gressions with the logarithm of response times for necessity not to not/lack of necessity interpretations as dv, verb type (mental/physical), context (necessity not to/lack of necessity) and their interaction as fixed effects, and item and participant as random effects (since the models with random slopes did not converge). we found no significant effects. we also ran the parallel analyses for likert rating times for necessity not to/lack of necessity interpretations, but also found no significant effects. as far as the variation group is concerned, we fitted the answers in the forced choice task into a generalized mixed effects model with answer as a dv, coded as 1 for lack of necessity readings and 0 for necessity not to readings, with verb type (mental/physical), context (necessity not to/lack of necessity) and their interaction as fixed effects and random slopes per item and participant. we found a significant effect of context (ß = 4.678, se = 1.174, z = 4.187, p < .01), but no other significant effect. a similar analysis with accuracy as a dv reveals no significant effect of context or other effects, showing that participants are quite sensitive to the context and verb type manipulations. in order to see the effect of context and/or verb type upon likert ratings, we fitted the data into a linear mixed effects model with likert rating as a dv and verb type, context and their interaction as fixed effects and random slopes per item and participant. we found a significant effect for the interaction between verb type and context (ß = −1.025, se = 0.446, df = 12.384, t = −2.294, p < .05). in necessity not to contexts, there was no significant difference in acceptability ratings for combining mustn’t with physical verbs (likert rating = 6.32) or mental verbs (likert rating = 6.15). however, in lack of necessity contexts, participants rated mustn’t in combination with physical verbs as 5.12, significantly lower than mustn’t in combination with mental verbs (6.38). in terms of response times in the forced choice task, we ran a mixed effects regression with the logarithm of response times for lack of necessity readings as a dv, verb type (mental/physical), context (necessity not to/lack of necessity) and their interaction as fixed effects, and random slopes per item and participant. we found no significant effect of context on the rts. we ran a parallel model for response times in the acceptability judgment task and found a significant effect for the interaction between context and verb type (ß = 0.853, se = 0.437, z = 66.9, t =1.940, p = .055). moreover, we also computed similar models for response times for necessity not to readings in the forced choice task and the acceptability judgment task, where we found no significant effects. 4. discussion. the interpretation of mustn’t in ae varies between speakers: for some ae speakers, mustn’t always expresses a necessity not to meaning, regardless of context, whereas for some ae speakers, mustn’t may express either a necessity not to meaning or a lack of necessity meaning, depending on pragmatic context and lexical verb type. interestingly, speakers who always preferred an interdiction reading for mustn’t tended to give lower ratings for the use of mustn’t in lack of necessity contexts, while speakers who were more context-sensitive also accepted mustn’t in both necessity not to and lack of necessity contexts with more ease. the variation observed for english is a regular pattern in other languages, e.g., romance languages where negation + obligation verb can contextually express either not necessary or necessary not to (see (8) for an example from romanian): (8) a. nu trebuie sǎ mergi la doctor. nu este un doctor bun. (romanian) not must să go to doctor. not be.3sg a doctor good ‘you must not go to the doctor. he is not a good doctor.’ b. nu trebuie sǎ mergi la doctor. eşti sǎnǎtoasǎ. not must să go to doctor. be.2sg healthy proceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 32 https://doi.org/10.3765/elm https://www.elm-conference.net/ ‘you need not go the doctor. you are healthy.’ several proposals may capture the availability of both the strong necessity not to reading and the weak lack of necessity reading in ae adult speakers. according to a neg-raising approach (hacquard 2010, homer 2011, iatridou & zeijlstra 2013) or an implicature account (jeretič 2021), the strong interdiction reading is derived from the basic order neg>modality via negative strengthening (modality>neg). according to pragmatic weakening (condoravdi 2012, von fintel 2019), the lack of necessity reading obtains as a suggestion from the basic strong necessity not to. according to an ambiguity approach, mustn’t is ambiguous between two basic readings (strong/weak). our results suggest that the two readings of mustn’t are not equally available to ae adult speakers, but rather mustn’t has a primary interdiction reading. please note that participants from the variation group interpret mustn’t as necessity not to to a certain degree even in lack of necessity contexts, and the participants who vary less in their interpretation (the interdiction group) understand mustn’t as expressing interdiction, not lack of necessity, independently of context. the fact that the necessity not to reading of mustn’t seems more available is in line with the pragmatic weakening account. moreover, the fact that participants rate physical verbs higher and faster than mental verbs in necessity not to contexts, but mental verbs higher and faster than mustn’t with physical verbs in lack of necessity contexts also supports an account based on sensitivity to pragmatic context and verb type. interestingly, we do not see this effect in response times in the forced choice task, but, as already mentioned, such data should be taken with a grain of salt. our results could also be accommodated by other accounts. the neg-raising approach could argue that negative strengthening is an automatic process which obtains by default from neg>modality in the case of mustn’t, such that its preferred reading is interdiction. negraising modals have to raise above negation due to their positive polarity sensitivity (homer 2011, 2015, iatridou & zeijlstra 2013). the weak reading would either be the primary, initial lf reading of mustn’t, or it would obtain by not doing or by cancelling negative strengthening when the context favors a lack of necessity reading. however, given the complex nature of negraising, we believe pragmatic weakening explains the results in a simpler, more natural fashion. to account for our results, the ambiguity approach would need to be modified such that, even though mustn’t has two available readings (a strong and a weak one), these are not equally accessible. more specifically, the interdiction reading is more salient, given that ae also makes use of needn’t, which is used exclusively to express a lack of necessity meaning. in line with grice’s manner maxim (1975) — see (9), we can assume that speakers prefer to avoid ambiguity, thus using mustn’t to primarily express the interdiction meaning: (9) maxim of manner: avoid obscurity; avoid ambiguity; be brief; be orderly. thus, our data can be best captured by pragmatic weakening and/or by an ambiguity approach which ascribes more salience to the interdiction reading of mustn’t instead of treating the two readings as equally available. in addition to context, the type of verb the modal combines with also matters. mental activities give rise to more lack of necessity readings than physical activities in lack of necessity contexts. necessity not to readings. our results can be explained within a cognitive account in terms of the difficulty of imposing one’s will over another’s (private) mental activities. from a speech act perspective, sentences containing mustn’t are “directives”, i.e., they convey “attempts to get the speaker to do something” (searle 1976:11). directives thus require that mustn’t should combine with a verb relating to doing something or a change-of-state verb. generally, proceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 33 https://doi.org/10.3765/elm https://www.elm-conference.net/ states do not occur in the imperative, since people cannot be ordered to be in a certain state (brown & miller 1980, bussmann 1996): statives “describe properties or relations which do not imply a change in state or motion and which cannot be directly controlled by the entity possessing the property, i.e. stative situations cannot be started, stopped, interrupted, or brought about easily or voluntarily” (busmann 1996:1120); in other words, statives are [-agentive] as the subject of the state is not consciously responsible for what happens (vendler 1957, dowty 1979, kearns 2000, griffiths 2006). the action restriction and the agentivity restriction account for the impossibility of directive speech acts with statives such as those in (10a) (such sentences would only be acceptable under an expressive wish reading), as well as the grammaticality of sentences with (durative) event verbs such as (10b). (10) a. *resemble your father! *have brown eyes! *understand the chaos theory! b. eat the pie! build a barn! walk in the park! interestingly, the mental verbs we used (worry, panic, be sad, be upset) are somewhat at the border between stative verbs and activity verbs. they are not action verb proper, but they can be conceived as dynamic if taken to express thought processes. they are also not agentive, but, rather, their subject expresses the thematic role of experiencer, being affected by an emotional or psychological state or experience (brown & miller 1980, kreider 1998, kearns 2000). according to crystal (2008), a clear classification of verbs into statives and dynamics is disturbed by the existence of such verbs which seem to evince properties belonging to both. there are indeed clear cases where statives can never be conceived as dynamic: one can never say *be tall, given that tallness cannot be envisaged as the outcome of an intentional act. however, other stative predicates are open to dynamic aspectual shifts: (11) be good! be quiet! don’t be stupid! know the answer by tomorrow! have a good time! although mental verbs such as worry, panic, be sad, be upset can be used in directives, it might still be perceived as somewhat odd to command someone not to engage in certain thought processes which are to a large extent out of the interlocutor’s control. hence, sentences with mustn’t and such mental verbs can be understood as expressing suggestions rather than interdictions. 5. conclusion in conclusion, in the current paper, we have provided experimental evidence that mustn’t may be understood differently by different ae adult speakers. while some speakers always interpret mustn’t as expressing interdiction, other speakers are sensitive to whether the context favors a lack of necessity or a necessity not to reading, or whether the verb the modal selects as a complement is a physical or a negative mental activity verb. essentially, these speakers tend to understand mustn’t as expressing a lack of necessity reading in lack of necessity contexts and with negative mental verbal complements. we have argued that our results support a pragmatic weakening account, where mustn’t has a primary interdiction reading, but, under certain conditions, speakers may weaken this interpretation. references coates, jennifer. 1983. the semantics of modal auxiliaries. london: croom helm. condoravdi, cleo & sven lauer. 2012. imperatives: meaning and illocutionary force. in christopher piñón (ed.), empirical issues in syntax and semantics 9, 37–58. http://www.cssp.cnrs.fr/eiss9/. proceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 34 https://doi.org/10.3765/elm https://www.elm-conference.net/ cormack, anna & neil, smith. 2002. modals and negation in english. in sjef barbiers, frits beukema, & wim van der wurff (eds.), modality and its interaction with the verbal system, 133–163. amsterdam: john benjamins. crystal, david. 2008. a dictionary of linguistics and phonetics. oxford: blackwell publishing. dowty, david. 1979. word meaning and montague grammar. dordrecht: reidel. grice, h. paul. 1975. logic and conversation. in peter cole, & jerry l. morgan (eds.), syntax and semantics, vol. 3, speech acts, 41-58. new york: academic press. griffiths, patrick. 2006. an introduction to english semantics and pragmatics. edinburgh: edinburgh university press ltd. homer, vincent. 2011. polarity and modality. los angeles, ca: university of california dissertation. fintel, kai von & sabine iatridou. 2019. a modest proposal for the meaning of imperatives. in ana arregui, maría luisa rivero, and andrés salanova (eds.), modality across syntactic categories, 288-319. oxford: oxford university press. hacquard, valentine. 2010. on the event relativity of modal auxiliaries. natural language semantics 18. 79–114. https://doi.org/10.1007/s11050-010-9056-4 haspelmath, martin. 1993. more on the typology of inchoative/causative verb alternations. in bernard comrie & maria polinsky (eds.), causatives and transitivity, 87–120. amsterdam: john benjamins. homer, vincent. 2015. neg-raising and positive polarity: the view from modals. semantics and pragmatics 8. https://doi.org/10.3765/sp.8.4. huddleston, rodney & geoffrey pullum. 2002. the cambridge grammar of the english language. cambridge: cambridge university press. 1859. iatridou, sabine & hedde zeijlstra. 2013. negation, polarity, and deontic modals. linguistic inquiry 44(4). 529–568. https://doi.org/10.1162/ling_a_00138. israel, michael. 1996. polarity sensitivity as lexical semantics. linguistics and philosophy 19, 619–666. https://doi.org/10.1007/bf00632710. jeretič, paloma. 2021. neg-raising modals and scaleless implicatures. new york: new york university dissertation. lingbuzz/006390. kearns, kate. 2000. modern linguistics semantics. new york: palgrave macmillan. kreidler, charles w. 1998. introducing english semantics. london: routledge. palmer, frank r. 1986. mood and modality. cambridge: cambridge university press. picallo, miguel. 1990. modal verbs in catalan. natural language and linguistic theory 8 (2). 285-312 papafragou, anna. 2000. modality: issues at the semantics-pragmatics interface. oxford, england: elsevier. r core team. 2021. r: a language and environment for statistical computing. r foundation for statistical computing. vienna, austria. available online at https://www.r-project.org/. ronai, eszter & ming xiang. 2020. pragmatic inferences are qud-sensitive: an experimental study. journal of linguistics 57(4), 841–870. https://doi.org/10.1017/s0022226720000389. searle, john r. 1976. the classification of illocutionary acts. language and society, 5. 1–23. vendler, zeno. 1957. verbs and times. the philosophical review, 66(2). 143–160. young, david. 1984. introducing english grammar. london: routledge. zejilstra, hedde. 2018. does neg-raising involve neg-raising? topoi 37(3). 417–433. proceedings of elm 2: 24-35, 2023 adina camelia bleotu, anton benz and roxana pǎtrunjel: the interpretation of mustn’t varies with context and verbal complement. 35 https://doi.org/10.3765/elm https://www.elm-conference.net/ reading times show effects of contextual complexity and uncertainty in comprehension of german universal quantifiers fabian schlotterbeck & petra augurzky* abstract. we report three experiments, in which we combined self-paced reading with picture-sentence verification to test how reading times are affected by meaning-related processes. in particular, we investigated german sentences containing the universal quantifier alle (“all”) and examined how restrictive processes incrementally interact with other aspects of quantifier meaning, comparably to previous studies using other methods. our results show that reading times were sensitive towards a match between context and sentence meaning and also towards an interaction between picture complexity and task demands. the results also point to the need for integrated processing models that combine refined notions of the relation between memory and expectations, on the one hand, with assumptions about adaptive processes and about representations involved in compositional interpretation, on the other. keywords. language comprehension; compositional-semantic processing; quantifier restriction; self-paced reading; expectation-based processing 1. introduction. current semantic and pragmatic theory offers detailed models of meaning-related processes for a wide range of linguistic phenomena. these models go beyond classical approaches in the sense that they not only intend to explain the compositional derivation of sentence meaning in general, but also focus on phenomena like incremental meaning composition (e.g. bott & sternefeld 2017), the complexity of meaning representations (e.g. pietroski et al. 2009, szymanik 2016) and contextual effects on the behavior of speakers and listeners (e.g. frank & goodman 2012, van tiel et al. 2021). despite these recent advances, relating predictions derived from semantic and pragmatic theory to processes during online comprehension remains an elusive goal in spite of the fact that theory-driven syntactic considerations have been implemented into models of on-line sentence comprehension for decades. this is especially surprising as highly comparable linking hypotheses could be developed on the basis of recent semantic and pragmatic models. for example, one could assume that complex meaning representations are generally avoided, or that highly expected sentence continuations lead to facilitation during incremental processing. we attempt to bridge this gap by studying how complexity and uncertainty in sentence meaning affect on-line sentence comprehension. in the current experiments, we combined self-paced reading with picture-sentence verification to test how reading times are affected by meaning-related processes. in particular, we examined how restrictive processes incrementally interact with other aspects of quantifier meaning, comparably to previous studies using other methods (augurzky et al. 2017, 2019, bott et al. 2019). *we would like to thank roman dick, hening wang & margarethe van liempt for help with programming and data preparation. authors: fabian schlotterbeck, university of tübingen (fabian.schlotterbeck@uni-tuebingen.de) & petra augurzky, goethe university frankfurt (augurzky@lingua.uni-frankfurt.de). fs received funding from the baden-württemberg ministry of science (mwk-bw) and the federal ministry of education and research (bmbf) as part of the excellence strategy of the german federal and state governments. pa received funding by project b1 of the dfg-funded sfb 833, and the dfg-/ahrc-funded project idealism. 1 proceedings of elm 2: 265-277, 2023 c©2023 fabian schlotterbeck and petra augurzky published by the lsa with permission of the author(s) under a cc by license. 265 https://doi.org/10.3765/elm https://www.elm-conference.net/ as an illustration of the basic phenomenon, consider the german sentences in (1), containing the universal quantifier alle (“all”), in the context of one of the pictures in fig. 1. (1) alle all dreiecke triangles sind are blau blue a. ..., ..., die that innerhalb inside .../ .../ b. ..., ..., die that außerhalb outside des of the kreises circle sind. are each picture showed colored objects, e.g. triangles, placed inside or outside a container shape, e.g. a circle. in that kind of setting, a truth-value judgment is, in principle, possible on the color adjective blue, but it may have to be revised later on if further restriction is provided, e.g. in the form of a restrictive relative clause (cf. example (1-a) in the context of fig. 1c where all and only the triangles inside the container shape are blue). previous studies following augurzky et al. (2017) repeatedly found evidence for an incremental truth evaluation on the adjective, but this effect was sensitive to the risk of a potential revision in the truth-value. in risky contexts, such as in fig 1c, in which a change in truth-value was possible, evidence for non-incremental truth evaluation was found. this finding was initially interpreted as an indication of revision-sensitivity: the processor plans ahead and only commits to an interpretation if there is no substantial risk to revise that interpretation later on. however, in subsequent studies, alternative accounts were also discussed, and we aim to compare these alternatives based on data from self-paced reading in the current study. in particular, we compare revision sensitivity to two competing accounts. the first alternative is based on pragmatic expectations. it assumes that comprehenders expect utterances that match the context, i.e. utterances that are true in the context (see e.g. augurzky et al. 2019, bott et al. 2019, schlotterbeck et al. 2022; for discussion). under this account, facilitation during processing is expected for words that allow for a high proportion of true vs. false continuations. the second alternative, which is also expectancy-based, explains modulation of processing difficulty in terms of priming. here the idea is that words denoting properties or concepts that are salient in the context lead to facilitation. note that all of these accounts are inherently interrelated. for example, revision sensitivity to some degree presupposes expectation-based processing (cf. schlotterbeck et al. 2022; for discussion) and pragmatic expectations may in part be driven by the salience of contextual properties (cf. frank & goodman 2012). despite these relations, we aimed to distinguish as clearly as possible between these three accounts in the following experiments. thus, we test a version of revision sensitivity that is not based on specific expectations but rather on the mere possibility of revisions of truth values. moreover, we take expectations that are based on salience rather than propositional meaning as evidence for priming rather than for pragmatic expectations. 2. experiment 1. the purpose of the first experiment was to replicate an erp experiment by augurzky et al. (2017) using self-paced reading. based on previous results, we expected to find facilitation for unambiguously true vs. false contexts but no such effects in risky contexts. the obtained results should serve as a baseline for further comparisons that are relevant to our three hypotheses, i.e. revision sensitivity, pragmatic expectations and priming. 2.1. methods. in each trial of the current self-paced reading experiment, participants first inspected a picture context showing geometrical objects inside and outside of a container shape (e.g. one of the figs. in 1a-d) and then read a universally quantified german sentence as in the examples in (1). half of the sentences contained a restrictive relative clause, as in (1-a/b), which could lead 2 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 266 https://doi.org/10.3765/elm https://www.elm-conference.net/ to a possible meaning change by introducing a subset reading (i.e. a reading where the quantificational domain is restricted to a subset of the geometrical objects, e.g. those outside the container shape). in total there were thus eight conditions in a 4 (context) × 2 (preposition) design. for sentences following the simple, unicolored contexts (fig. 1a/b), a truth-value judgment is possible already on the adjective (blau, ‘blue’). by contrast, for bicolored contexts (fig 1c/d), the judgment has to be delayed until the preposition (innerhalb, ‘inside-of’, außerhalb, ‘outside-of’) is encountered. eight items per condition were constructed by combining various color adjectives (blue, red; green, yellow; orange, purple; gray, black), geometric shapes (triangles, pentagons, semicircles, hearts) and container shapes (rectangles, circles), yielding 64 experimental sentences. these were matched with 64 short filler sentences (as in (1), each presented twice). in addition, there were 16 catch trials, which prompted participants to make truth-value judgments about simple logical statements presented without contexts. in total this resulted in 144 trials that were distributed over four blocks using a latin square design. sentences were presented word-by-word (with punctuation displayed separately) using the moving-window technique, and participants performed a truth-value judgment task after each sentence. fifty-nine german native speakers were recruited over the platform prolific.co and were paid for their participation. we excluded five participants from further analysis because they met one of the following exclusion criteria, which we used in all three experiments: word reading times were over 10 s in at least one trial; overall accuracy was below 70%; performance on catch trials was below 70%; or duration of the entire experiment was more than 3 standard deviations above the mean. 2.1.1. statistical analysis. for inferential statistics we used mixed-effects models (bates et al. 2015) that were similar in the current and the following two experiments, with only a few differences due to adaptations to each experimental design. for accuracy, we used logit mixedeffects models, which included the fixed effects of context, preposition and their interaction. rts were trimmed individually for each participant and condition by removing extreme rts that were shorter than 200 ms or longer than three standard deviations above the participant’s mean in that condition. for rts on the adjective, linear mixed-effects models were computed that predicted log-transformed rts based on context. for each of the experiments, custom contrasts for the multilevel factor context were specified to test our hypotheses (see table 1). they are explained in the results sections of the individual experiments. for rts on the preposition, the factor preposition was also included, but for simplicity we focus on effects of truth values (i.e. specific interactions between context and preposition) as they were predicted by revision sensitivity. other effects are only mentioned if theoretically relevant. for some effects we had derived directed predictions from our hypotheses. we thus calculated one-tailed t-tests for them and marked them as such in the results sections. in all models, by-participant and by-item random intercepts and slopes were included if they allowed for convergence and contributed to model fit. 1 2.2. results. across conditions, truth-value judgments were correct in the majority of cases (94.0% − 98.7%; see fig. 2b). the highest accuracy was achieved in the false unicolored conditions, leading to a marginal effect of context (χ2(3) = 6.34, p = .096), but, at the same time, no significant difference to any of the other context conditions was found (|z| < 1.6). reading 1in the third experiment below, only random intercepts were included for participants due to the between design. 3 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 267 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) one color, match (b) one color, mismatch (c) two colors, match inside (d) two colors, match outside (e) three colors, match inside (f) five colors match inside (g) three colors, mismatch inside (h) five colors, mismatch inside (i) three colors, match outside (j) five colors, match outside (k) three colors, mismatch outside (l) five colors, mismatch outside figure 1: sample visual contexts (exp. 1: a-d; exp. 2: a-d, e & i ; exp. 3: a-l) exp .1 exp .2 exp. 3 # contrast # contrast # contrast # contd. contrast # contd. contrast 1 a, b vs c, d 1 a, b vs c-e, i 1 a, b vs c–l 6 e, f vs. g, h 9 i, j vs. k, l 2 a vs. b 2 a vs. b 2 a vs. b 7 g vs. h 10 k vs. l 3 c vs. d 3 c vs. d 3 c vs. d 8 e vs. f 11 i vs. j 4 c, d vs. e, i 4 c, d vs. e–l 5 e vs. i 5 e–h vs. i–l table 1: contrasts for the factor context used in the three experiments: labels of subfigures in fig. 1 are used as short references to context conditions times at the adjective are shown in fig. 2a. for the unicolored contexts, true conditions were read faster than false ones (one-tailed: t = 1.66, p = .049). for the bicolored contexts, a marginal difference between context where the matching color was inside vs. outside the container shape was found (t = 1.7, p = .086). moreover, bicolored contexts led to significantly longer rt than unicolored ones (t = 2.4, p = .017). the latter effect was sustained over several regions and turned out to be reliable on the preposition as well (β = 0.11, t = 5.81, p < .001). moreover, there was a truth-value effect for the bicolored contexts on the preposition (context×preposition: 4 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 268 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) mean rt on adjective in exp. 1 (b) accuracy in exp. 1 figure 2: results from exp. 1 χ2(1) = 4.07, p = .044) which was due to shorter rt for the matching than mismatching bicolored contexts (significant for the preposition inside: means (sds): 332 (184) vs. 374 (299), t = 2.7, p = .007; but not outside: means (sds): 355 (184) vs. 350 (299), t = −0.15, p = .007). 2.3. discussion. the current self-paced reading results indicate that, comparable to erps (augurzky et al. 2017), rts are generally sensitive to the truth value of an utterance: we observed effects of truth values as soon as they could be unambiguously decided, i.e. on the color adjective in unicolored contexts and on the preposition in bicolored contexts. these effects may reflect a local truth evaluation and are thus in accord with revision sensitivity. however, they are also in line with pragmatic expectations as they may alternatively signal a facilitation for expected sentence continuations that allow for true descriptions of the context. finally, these results are also compatible with priming, as they could reflect facilitation effects for reading words denoting properties that were salient in the visual contexts. however, purely lexical priming cannot explain the truth-value effect at the preposition, as the truth value at this point depends on both the preposition and the adjective. thus, the salience of the matching preposition in the visual context cannot be the sole source of facilitation. we may instead assume that combined properties are primed (e.g. blue-inside), implying a rudimentary from of composition, or that the adjective is a cue for memory retrieval of the matching preposition. under both of these assumptions, the hypothesis of priming would look a bit more similar to pragmatic expectations. the following two experiments were aimed to distinguish between these competing explanations. on top of these potentially truth-value related effects, we observed a substantial slowdown for the bivs. unicolored picture contexts across several words. we interpret these effects as being related to memory demands: whereas only one color term had to be remembered in the unicolored contexts, two colors and their positions were relevant to the experimental task in the bicolored contexts. below, in the general discussion, we reflect on theoretical explanations of this effect. for now it is simply important to acknowledge this effect while we attempt to dissect the truth-value 5 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 269 https://doi.org/10.3765/elm https://www.elm-conference.net/ related effects in the following two experiments. 3. experiment 2: testing for expectation-based effects. in order to disentangle the two expectation-based accounts (i.e. priming and pragmatic expectations) from revision sensitivity, we included another pair of visual contexts into the experimental design (see fig. 1e/i). in these pictures, one element of the mismatching set of objects (inside or outside the container shape) had a different, third color that was not mentioned in the sentence. this change does not affect the predictions from revision sensitivity: the newly added contexts still make the sentence false locally on the adjective and are also still risky in the sense that our test sentences may turn out to be true depending on the specific continuation (e.g. ... that are inside of the circle). the situation is different under the two expectation-based accounts. firstly, under pragmatic expectations, predictions differ between the newly added contexts as compared to the original complex contexts in fig. 1c/d. this is because in the former, there is only one true continuation among the presented alternatives. in this sense, these types of complex contexts in fig. 1e/i are comparable to the simple contexts in fig. 1a/b. therefore, the prediction from pragmatic expectations is that the two newly added contexts should lead to faster rts than the complex contexts in exp. 1 because the adjective (e.g. blue) is more expected by virtue of allowing for true relative clause continuations. secondly, under priming, predictions depend on how visual salience is operationalized. if we think of salience of a color in terms of the fraction of the total area it covers, then predictions remain unchanged. if we instead think of it in terms of number of different colors shown in the picture, salience of the mentioned color (e.g. blue) will, by contrast, be altered, and presumably reduced, in the two new contexts because there is a third color beside the original two. moreover, we also have to consider the complexity effect in univs. bicolored contexts in exp. 1. this complicates predictions because at the moment we are lacking a theoretical understanding of this effect. therefore, we do not know how it will be affected by the additional color in the newly added contexts and whether interacts with other factors affecting processing difficulty. 3.1. methods. the present experiment used the same general method and procedure as exp. 1. the experimental design and materials were adjusted slightly by including additional contexts. the materials from exp. 1 were reused and combined with the two new types of contexts exemplified in fig. 1e/i. this resulted in twelve conditions in a 6×2 design with the factors context (e.g. fig. 1a–e,i) and preposition (invs. outside). in total, eight items per condition were constructed by combining the same colors and geometrical shapes as in exp. 1, yielding 96 experimental sentences that were matched with 96 short filler sentences (each presented twice). combined with 32 catch trials of the same type as in exp. 1, this resulted in 224 trials that were distributed over four blocks using a latin square design. fifty native speakers of german who hadn’t participated in exp. 1 were recruited over the platform prolific.co and were paid for their participation. 2 the same exclusion criteria as in exp. 1 resulted in the exclusion of four participants. 3.2. results. accuracy across conditions is shown in fig 3b. as in exp. 1 participants’ responses were correct in the majority of cases (92.3 − 97.7%). accuracy was lower and showed more variation in the biand tricolored as compared to the unicolored contexts. the logit mixed2there were 58 participants originally, but 8 of them had to be excluded because of an error with data transfer to the experimental server resulting in rt that could not be unambiguosly assigned to roi. 6 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 270 https://doi.org/10.3765/elm https://www.elm-conference.net/ (a) mean rt on adjective in exp. 2 (b) accuracy in exp. 2 figure 3: results from exp. 2 effects model analysis (see section 2.1.1) revealed a marginally significant interaction between the two factors (χ2(5) = 9.65, p < .086; based on model comparison) showing that, in the two newly added conditions, the position of the homogeneous, i.e unicolored, color set had numerically opposite effects depending on the preposition in the sentence. when the homogeneous set was inside the container shape, the preposition outside led to slightly higher accuracy. by contrast, when the homogeneous set was outside the container shape, it led to slightly lower accuracy. in separate analyses of the two new contexts, we found that the interaction between context and preposition was significant (z = 2.27, p < .023). however, the differences between prepositions was not reliable for either position of the homogeneous set (match inside: z = −0.29, p = .766; match outside: z = 1.69, p = .091). within the other contexts, no effects were significant (all |z| < 1.6). rts on the adjective are shown in fig. 3a. the general pattern in the uniand bicolored contexts replicates the results from exp. 1. moreover, tricolored contexts had comparable rts to bicolored contexts. in order to test our hypotheses, we computed a mixed model including the five contrasts in the second column of table 1 for the six-level factor context. based on exp. 1, we predicted that unicolored and locally true contexts would lead to shorter rts than complex (biand tricolored) and locally false contexts, respectively (contrasts 1 & 2 in table 1). both predictions were borne out (simple vs. complex: t = 6.75, p < .001; locally true vs. false, one-tailed: t = 1.81, p = .038). for the remaining three contrasts, no directed prediction could be derived from our hypotheses, especially due to uncertainty regarding the contribution of the complexity effect, and thus we computed two-tailed hypothesis tests. we found that adjectives were read marginally faster after tricolored vs. bicolored contexts (contrast 3: t = 1.93, p = .059), but the location of the homogeneous color set did not affect rt significantly (t = 0.93, in bicolored contexts; t = 1.58, in tricolored contexts; contrasts 4 and 5, respectively). turning to the preposition, model comparisons revealed that, comparable to the results at the adjective, the preposition was read faster after unicolored than after bicolored contexts (t = 6.52, p < .001). furthermore, the unicolored false contexts were read faster than the unicolored true ones (t = −3.57, p < .001). no other effects were significant on the preposition. 7 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 271 https://doi.org/10.3765/elm https://www.elm-conference.net/ 3.3. discussion. although the newly added contexts, containing one object in a different color, were presumably at least as complex to remember as the bicolored contexts from exp. 1, they did not lead to an increased complexity effect in rt and were even read faster on the color adjective, leading to marginal effect in the opposite direction. assuming that tricolored contexts were more taxing on memory than bicolored ones, rt on the adjective thus seem to reflect a speed-up in anticipatory processes after tricolored contexts. under pragmatic expectations, such a speedup in anticipatory processing could be due to the fact that only one of the possible continuations may yield a true output, e.g., blue that are inside the circle and is thus expected more strongly in trivs. bicolored contexts. at the same time, remembering three colors and their positions is more taxing on memory than remembering only two. thus, total rt, reflecting both anticipatory and memory-related processes, may end up in a range similar to that of the bicolored contexts. under priming, anticipatory processes on the color adjective may have been sped-up because of an increased salience of the relevant color term in the tricolored contexts. in these contexts, the mentioned color not only had the most representatives but also formed the only homogeneous set, i.e. covered all the object in its region. this may by itself induce an expectation that this is the color that will be referred to, again increasing expectations and resulting in a relative speed-up that may possibly be overlayed with an increased complexity effect. a salience-based explanation along these lines could, however, be questioned on the basis of the accuracy data. if the interaction between preposition and context in the tricolored contexts is real, it would point to the additional object directing attention away rather than towards the homogeneous set. if this was in fact the case, increased salience of the color of the homogeneous set would be implausible. finally, under revision sensitivity we would have expected no differences in rt between multicolored contexts, apart from effects related to visual complexity and memory. this is because, in all these conditions, the sentence is false locally on the color adjective, but risky in the sense that it may still end up true given the right continuation. from this perspective the newly added, tricolored contexts were expected to behave exactly as the bicolored. combined with their increased visual complexity we would have expected an increase rather than decrease in reading times on the adjective. 4. experiment 3: contrasting expectationand memory-based effects. to disentangle effects that are based on memory demands from effects based on anticipatory processing as best as possible, we included an additional set of visual contexts in the present experiment which consisted of variations of the tricolored contexts in exp. 2. some of these contexts showed even more colors in the heterogeneous set. according to priming, additional colors should affect the salience of color terms and lead to a decrease in rts. these additional colors are, however, not expected to affect pragmatic expectations beyond what salience contributes. this is because we assumed pragmatic expectations to be driven by an expectation of true utterances (or utterances that match the context). moreover, from the perspective of revision sensitivity adding more colors to the heterogeneous set in contexts such as fig. 1e/i does not change anything. sentences in these conditions are still locally true but risky. these new contexts should, therefore, be processed just as the multicolored contexts in exps. 1 & 2, modulo potential effects of visual complexity. in addition, we created a set of contexts in which we also changed the color of the homogeneous set (e.g. from blue to red, i.e. from matching to mismatching). this is expected to lead to a slowdown 8 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 272 https://doi.org/10.3765/elm https://www.elm-conference.net/ under both priming and pragmatic expectations, either because of the increased salience of the mismatching color or because of the unexpectedness of an adjective that does not allow for true continuations (in the present design). regarding memory demands, we would expect that contexts with more colors pose higher burdens on memory and thus lead to longer rt. the reasoning behind this expectation is that the more colors are shown on the picture the more information there is to be remembered, in particular the color of objects and their positions. 4.1. methods. the same materials as in exp. 2 were used and these were combined with six further types of multicolored contexts, exemplified in fig. 1f–h,j–l. within the eight multicolored contexts that were not included in exp. 1 (fig. 1e–1l), three factors were manipulated in a 2 (number of colors: 3 vs. 5) × 2 (homogeneous color: match vs mismatch) × 2 (position of homogeneuous set: inside vs. outside) factorial design. in combination with the original four contexts, this resulted in 24 conditions in a 12 (context: see fig. 1a–l) × 2 (preposition: inside vs. outside) within-items design. the entire set of visual contexts allows for a number of comparisons relevant to our three hypotheses and to the nature of the complexity effect from exps. 1 & 2. these comparisons are described below in the results section. to keep the number of trials per participant in an acceptable range, especially for an online experiment, conditions were distributed between participants in eight lists. in each list, the four contexts from exp. 1 were combined with two of the trior quintcolored contexts. each context condition was included in two of the eight lists and that triand quintcolored contexts were never combined in one list. context conditions were distributed across lists according to the following scheme (using labels from fig. 1 to refer to conditions: list 1: e & i; list 2: f & j; list 3: f & l; list 4: h & l; list 5: h & j; list 6: e & k; list 7: g & k; list 8: g & i). the lists were generated completely analogously to exp 2. and one of them was, in fact, identical to the materials in exp. 2. in total, 281 native speakers of german (between 34 and 37 per list) who hadn’t participated in one of the previous experiments or in any of the other lists were recruited over the platform prolific.co and were paid for their participation. the same exclusion criteria as before were used and resulted in exclusion of 53 participants in total from further analysis. 4.2. results. while accuracy was generally high (92.6%−98.7%, see fig. 4b), the highest values were observed in quintcolored contexts. for simplicity, we restricted the inferential statistics of accuracy to the triand quintcolored contexts. a logit mixed effects model analysis with fixed effects of number of colors, homogeneous color, position of homogeneous set and preposition as well as random intercepts of participants revealed two significant interactions (position of homogeneous set × preposition: z = 2.4, p = .019; and number of colors× homogeneous color: z = −4.18, p < .001). we resolved these interactions in separate analyses of triand quintcolored contexts. in fact, the pattern differed markedly between these two subsets: in the tricolored contexts, the general pattern from exp. 2 was replicated. thus, the preposition outside led to fewer errors if the homogeneous set was inside and vice versa. although this pattern was descriptively visible in each of the tricolored contexts, it was only significant in conditions with mismatching contexts and preposition inside (z = 2.69, p = .007), leading to a significant three-way interaction in the tricolored contexts (z = 2.45, p = .015). by contrast, only a main effect of homogeneous color (z = −4.89, p < .001) was found in the quintcolored contexts because mismatching colors led to more correct responses than matching ones. 9 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 273 https://doi.org/10.3765/elm https://www.elm-conference.net/ 270 290 310 330 350 1col, tr ue 1col, fa lse 2col, m atch inside 2col, m atch outside 3col, m atch inside 5col, m atch inside 3col, m ismatch inside 5col, m ismatch inside 3col, m atch outside 5col, m atch outside 3col, m ismatch outside 5col, m ismatch outside context m ea n r t ( an d 95 % c on fid en ce in te rv al s) predicted values of rt (a) predicted rt at the adjective (b) accurcy in exp. 3 figure 4: results from exp. 3 the between-subjects design led to considerable variation in rt within the triand quintcolored contexts. to adjust for differences between participants, we provide a plot of predicted rt (i.e. marginal effects) from the mixed model analysis (fig 4a). the mixed-effects model of rt at the adjective (see section 2.1.1) included comparisons analogous to those in exps. 1 & 2 (i.e., contrasts 1–5 in the third column of table 1) in addition to six further comparisons (contrasts 6–11 in table 1) that are restricted to triand quintcolored contexts and tested for specific predictions of priming and pragmatic expectations. the contrasts 6/9 tested for a prediction of pragmatic expectations (in contexts with the homogeneous set inside or outside the container shape, respectively) that adjectives like blue should be read faster if they start the only true continuation (e.g. 1e/f) than if a different adjective starts the only such continuation (e.g. 1g/h). the contrasts in 7/10 and 8/11 tested for the prediction of priming (with the homogeneous set inside or outside, respectively) that a relatively large number of objects in a distracting color in a matching contexts (e.g. fig 1e vs.f and 1i vs. j) or a relatively small number of matching objects in a mismatching contexts (e.g. fig 1g vs. h and 1k vs. l) is expected to cause slow down. the analysis revealed significant effects of the contrasts 1, indicating shorter rt in unicolored contexts in comparison to the mean rt of all the others (t = 21.63, p < .001), 2, indicating shorter rt for true vs. false unicolored contexts (t = 4.57, p < .001), and 3, showing longer rt after bicolored than after triand quintcolored contexts (t = 7.286, p < .001). these three effects were replicated from exp. 2 (and partly exp. 1). moreover, contrast 8, comparing the number of mismatching (e.g. red) objects in three vs. five colors, match inside (t = −5.01, p < .001), and contrast 10, comparing the number of matching objects in three vs. five colors, mismatch outside (one-tailed: t = −2.72, p = .003), were significant. and finally, the contrast 7 was significant (t = −2.48, p = .013) showing that more objects in a matching color led to a slow down in mismatch inside contexts. none of the other contrast were significant (|t| < 1.3). looking at the direction of the effects, only contrast 8 seems to speak to one of our hypotheses, namely priming, because it seems to show that a larger number of objects in a mismatching color slows 10 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 274 https://doi.org/10.3765/elm https://www.elm-conference.net/ down reading. however, all of the effects in the contexts with three and five colors could also be explained by a general trend for faster rt in contexts with five vs. three colors in addition to an interaction between homogeneous color and position of homogeneous set in the tricolored contexts. to test for this explanation, we opted for a post-hoc, factorial analysis of the triand quintcolored contexts (cf. description of factors above). this analysis revealed a three-way interaction in (t = 2.52, p = .012), which we resolved by conducting separate analysis of the triand quintcolored contexts. there was a two-way interaction in the former (t = −2.52, p = .012) but not the latter (t = 0.17, p = .87). the interaction in the tricolored contexts was due to a marginally significant difference between the two matching (t = −1.68, p = .094) and an opposite but non-significant difference in the mismatching contexts (t = 1.09, p = .276). on the preposition, we observed an intricate pattern of rt. because there was no truth-value effect, we refrain from reporting it in detail. one finding that seems worth mentioning, however, was a general tendency for longer rt in matching sentences (i.e., sentences that were potentially true at the adjective) vs. mismatching sentences. in multicolored contexts with the homogeneous set inside the container shape, this was the only significant effect (t = 2.8, p = .005). when the homogeneous set was outside, it interacted with the number of colors but was significant within both levels of this factor (3 cols.: t = 4.24, p < .001; 5 cols.: t = 8.9, p < .001). 4.3. discussion. while pragmatic expectations can provide a plausible explanation of the results from exps. 1 & 2 (at least in combination with certain assumptions about memory-related effects), its predictions were disconfirmed in the present experiment. in particular, rts did not differ between contexts that allowed for true and false continuations. moreover, priming only correctly predicted the difference between contexts with three and five colors. in general, contexts with more colors actually led to shorter rt, contrary to our initial expectations. one possible explanation could be that the more colors that were present, the better participants were able to focus on the homogeneuos set and ignore the distractor objects. this would actually be a useful adaptation to the present design, because only the homogeneous sets were relevant for the task. taking into account this reversed complexity effect, the data on the adjective seem compatible with revision sensitivity. however, we did not find an effect of truth on the preposition. instead, we saw that the conditions that were potentially true on the color adjective led to prolonged rts on the preposition, regardless of their actual truth values at that point. we take this effect as an indication of memory retrieval of the position of the matching color set. if this interpretation is on the right track, our entire data set would call for an explanation in which truth-evaluation interacts in intricate ways with effects that are related to the retention of information in and retrieval of that information from memory. this interpretation also seems to be consistent with the interactions we observed in the tricolored contexts in rt on the adjective and also in accuracy. these interactions indicate that the one objects that had a different color in these conditions may modulate memory load by attracting attention to varying degrees depending on the position of the homogeneous set. 5. general discussion. overall, the present experiments show that self-paced reading is affected by the relation between a previously presented visual context and the incremental processing of the compositional meaning of a quantified sentence. in particular, reading times were sensitive towards a match between context and sentence meaning (e.g. the truth-value related effects in exp. 1) and also towards interactions between picture complexity and task demands (e.g. the difference 11 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 275 https://doi.org/10.3765/elm https://www.elm-conference.net/ between uniand multicolored pictures and its modulation across the three experiments). we also observed rt effects that seemed to depend on a combination of these two factors (e.g. the effect of matching vs. mismatching colors at the preposition in exp. 3). crucially, none of our three hypotheses can account for our entire data set. first of all, revision sensitivity predicted truth-value effects as soon as they can be unambiguously decided, but, on the preposition, such effects were only observed in exp. 1. secondly, pragmatic expectations predicted rt effects already on the adjective, which is the position where the potential for true continuations is decided. this was directly tested in exp. 3 but not confirmed. finally, priming predicted that the number of objects that match or mismatch the color adjective should affect rts, irrespective of the (potential) truth value, which was also not confirmed in exp. 3. in the following, we suggest two possible explanations for this discrepancy between predictions and results. firstly, some of the predicted effects could have remained undetected because of insufficient power. in particular, we consider it possible that the complexity effects may have obscured a potential truth-value effect. this appears likely if we consider the relatively large magnitude (and associated variance) of the complexity effect compared to the truth-value effect that we consistently observed within the unicolored contexts. however, note that we did not have any specific prior expectations about these effects. another, factor that may have limited the chance to detect truth-value related effects could be participant variance caused by the between-design in exp. 3. this points to a general methodological challenge in distinguishing between the three tested hypotheses: because of the close relation between these hypotheses, distinguishing between them requires specific comparisons and the resulting number of relevant conditions may exceeded the limits of within-designs. one way to approach this issue would be to use continuous independent variables to manipulate pragmatic expectations and priming in a gradual fashion. secondly, the discrepancy between predictions and results may also be related to working memory limitations. specifically, the task of encoding all information in the context pictures and using this information to build expectations about all possible continuations in our task is cognitively highly demanding. indeed, there are theoretical considerations on effects of restricting the memory capacity underlying pragmatic interpretation (franke et al. 2011) as well as proposals that make expectations during sentence processing contingent on imperfect memory representations (e.g. futrell et al. 2020) or on adaptations to the experimental context (see e.g. schlotterbeck et al. 2022; for an application to quantifier processing). if the observed discrepancy is due to the second, theoretically more interesting explanation, we would still be missing crucial pieces to the puzzle of incremental compositional interpretation. we think that applying refined notions of the relation between memory and expectations to semantic and pragmatic processing may lead us a crucial step forward towards finding missing pieces. moreover, such considerations can be combined with assumptions about adaptive processes as well as specific theories about the representations and processes involved in compositional interpretation (see e.g. bott et al. 2019, bremnes et al. 2022, pietroski et al. 2009; for application to quantifiers) in order to refine the type of hypotheses tested in the present study. if such hypotheses are tested using a mix of different methods, we may obtain a better understanding of the processing correlates of compositional interpretation. 12 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 276 https://doi.org/10.3765/elm https://www.elm-conference.net/ references. augurzky, p., o. bott, w. sternefeld & r. ulrich. 2017. are all the triangles blue?–erp evidence for the incremental processing of german quantifier restriction. language and cognition 9(4). 603–636. https://doi.org/10.1016/j.cognition.2017.05.023. augurzky, petra, michael franke & rolf ulrich. 2019. gricean expectations in online sentence comprehension: an erp study on the processing of scalar inferences. cognitive sci 43. https://doi.org/10.1111/cogs.12776. bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67(1). 1–48. https://doi.org/10.18637/jss.v067.i01. bott, oliver, fabian schlotterbeck & udo klein. 2019. empty-set effects in quantifier interpretation. journal of semantics 36(1). 99–163. https://doi.org/10.1093/jos/ffy015. bott, oliver & wolfgang sternefeld. 2017. an event semantics with continuations for incremental interpretation. journal of semantics 34(2). 201–236. https://doi.org/10.1093/jos/ffw013. bremnes, h.s., j. szymanik & g. baggio. 2022. computational complexity explains neural differences in quantifier verification. cognition https://doi.org/10.1016/j.cognition.2022.105013. frank, m. c. & n. d. goodman. 2012. predicting pragmatic reasoning in language games. science 336(6084). 998–998. https://doi.org/10.1126/science.1218633. franke, michael, gerhard jäger & robert van rooij. 2011. vagueness, signaling and bounded rationality. in takashi onada, daisuke bekki & elin mccready (eds.), new frontiers in artificial intelligence, 45–59. berlin, heidelberg: springer berlin heidelberg. https://doi.org/10.1007/978-3-642-25655-4 5. futrell, richard, edward gibson & roger p. levy. 2020. lossy-context surprisal: an information-theoretic model of memory effects in sentence processing. cognitive science 44(3). https://doi.org/10.1111/cogs.12814. pietroski, paul, jeffrey lidz, tim hunter & justin halberda. 2009. the meaning of ‘most’: semantics, numerosity, and psychology. mind and language 24(5). 554–585. https://doi.org/10.1111/j.1468-0017.2009.01374.x. schlotterbeck, fabian, petra augurzky & rolf ulrich. 2022. degree of incrementality is modulated by experimental context – erp evidence from german quantifier restriction. language, cognition and neuroscience https://doi.org/10.1080/23273798.2022.2125541. szymanik, jakub. 2016. quantifiers and cognition. logical and computational perspectives. springer. http://doi.org/10.1007/978-3-319-28749-2. van tiel, bob, michael franke & uli sauerland. 2021. probabilistic pragmatics explains gradience and focality in natural language quantification. proceedings of the national academy of sciences of the united states of america 118(9). https://doi.org/10.1073/pnas.2005453118. 13 proceedings of elm 2: 265-277, 2023 fabian schlotterbeck and petra augurzky: contextual complexity and uncertainty in comprehension of german universal quantifiers. 277 https://doi.org/10.3765/elm https://www.elm-conference.net/ testing the influence of quds on the occurrence of conditional perfection britta grusdt, mingya liu, & michael franke* abstract. in natural language conversations, speakers often communicate ‘if and only if’ when they say ‘if’. the reasons why in some circumstances, yet not all, conditionals receive a biconditional interpretation remain under investigation. von fintel (2001) proposed an account where the interpretation of a conditional (“if p, then q”) is predicted to depend on the focus of the conversation which may either lie on the conditions that make the consequent, q, true or on the consequences following when the antecedent, p, is true. to test this account, we present two novel behavioral experiments with non-text based stimuli that take advantage of participants’ intuitive understanding of physics. we find some supporting evidence for the tested account that is not conclusive but suggests that other aspects, like the nature of potential alternative causes for the consequent to become true (e.g., with or w/o the influence of an external variable), also play a role for the interpretation of the conditional. keywords. conditional perfection; indicative conditionals; pragmatics; question under discussion; exhaustivity; experimental 1. introduction. a longstanding subject of research in the context of indicative conditionals — natural language expressions of the form “if p, q” (p→ q) where p is referred to as antecedent and q as consequent— is their interpretation as biconditionals, a phenomenon known as conditional perfection (cp) (geis & zwicky 1971). a conditional is said to be perfected when ‘if’ communicates ‘if and only if’, leaving unaddressed how exactly this inference comes about, e.g. through one of the inferences ‘only if p, q’, ‘if not p, not q’ or ‘if q, p’ (see van canegem-ardijns & van belle (2008) who propose different types of conditional perfection). for some conditionals, a perfected interpretation seems to be readily endorsed, even when no concrete context is given; an eminent example are promises and threats communicated with conditionals like (1) below. here, the biconditional interpretation is forthcoming naturally: the speaker seems to communicate to scratch the addressee’s back if and only if the addressee scratches the speaker’s back. (1) if you scratch my back, i’ll scratch yours. in the literature on conditional reasoning, participants are usually tested on four inferences, denying the antecedent (da), affirming the consequent (ac), modus ponens (mp) and modus tollens (mt) to investigate their interpretation of conditionals.1 since, contrary to mp and mt, da and ac are only logically valid in the case of biconditionals, high endorsement rates of the latter two inferences suggest a biconditional interpretation; fillenbaum (1986) reports endorsement rates of da for conditional premises and threats between 80 and 90%, whereas they tend to *we would like to thank malin spaniol and josefine zerbe for their great support in implementing the experiments. authors: britta grusdt, universität osnabrück (britta.grusdt@uni-osnabrueck.de) & mingya liu, humboldt universität zu berlin (mingya.liu@hu-berlin.de) & michael franke, eberhard karls universität tübingen (michael.franke@uni-tuebingen.de). 1tabular overview of the four classical conditional reasoning inferences: . . . continues on next page . . . proceedings of elm 2: 104-116, 2023 c©2023 britta grusdt, mingya liu and michael franke published by the lsa with permission of the author(s) under a cc by license. 104 https://doi.org/10.3765/elm https://www.elm-conference.net/ be (much) lower for other types of conditionals (e.g. see evans et al. 1993; chapter 2 for a summary of studies). another factor that has been shown to influence participants’ interpretation of a conditional as biconditional is the availability of alternative causes or disabling conditions for the consequent, making a biconditional interpretation more likely when fewer alternative causes are conceivable (cummins et al. 1991, markovits 1986). this finding connects with the hypothesis of how cp arises that we aim to test in this paper, as we will see shortly. in the linguistic literature, cp-readings of conditionals have often been treated as conversational implicatures (atlas & levinson 1981, van der auwera 1997, horn 2000). while geis & zwicky, who (re)initiated the debate about cp among linguists,2 argue that conditionals are quite commonly attributed a cp-reading, this regularity has been questioned by others providing various counterexamples (e.g., lilje 1972, de cornulier 1983). von fintel (2001) goes one step further by making a proposal of when exactly a conditional is interpreted as biconditional and when it is not, which had not been precisely formulated in previous accounts. similar to de cornulier (1983), von fintel refers to exhaustivity: he argues that a biconditional interpretation arises when the antecedent of a conditional is interpreted as exhaustive list of conditions for the consequent whereas it is not triggered when the conditional is interpreted as exhaustive list of consequences of the antecedent. in other words, when the speaker is required to provide an exhaustive list of conditions for the consequent (q) and only mentions a single condition (p) the listener will infer that p is a sufficient and necessary condition for q, corresponding to a cp-reading. on the other hand, mentioning p as single condition in the antecedent does not trigger a cp-reading when the speaker is asked to provide an (exhaustive) list of consequences of p. according to von fintel (2001), a (possibly implicit) question under discussion (qud) determines whether the conditional targets the conditions of the consequent (e.g., under which conditions q?) or the consequences of the antecedent (e.g., what follows from p?). to illustrate the hypothesized effect of the qud, consider the conditional in (2), inspired by an example from lilje (1972): (2) if a ball bounces off the table, it is a foul. in a situation where a person a explains the rules of the game pool to a person b who has no experience with this game and where b asks a ... (i) what happens if a ball bounces off the table (qud: if-p) (ii) which actions count as foul / whether there are actions that count as foul (qud: when-q) the same answer — the conditional in (2) — seems to be interpreted as biconditional only when the conversation is guided by the question in (ii). given the context provided by the qud if-p (i), the speaker is not expected to mention all possible actions that are considered a foul and, thus, cp does not arise in this case. premise 1 premise 2 conclusion inference logically valid p→ q. p. ∴ q. modus ponens (mp) 3 p→ q. ¬ q. ∴ ¬ p. modus tollens (mt) 3 p→ q. q. ∴ p. affirming the consequent (ac) 7 p→ q. ¬p. ∴ ¬ q. denying the antecedent (da) 7 2as noted by van der auwera (1997), cp had already been discussed before geis & zwicky (1971), e.g. by ducrot (1969). proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 105 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 1: exhaustive and non-exhaustive situations of critical trials in experiment 2. the scenes in experiment 1 were identical except for the yellow distractor block in stimuli c and d which was also centered on the platform, but standing on its short side. for simplicity, in all pictures shown here, the ant-block is green and thus, the cons-block blue; in the experiments, the color of the antecedentand the cons-block was randomly chosen for each participant and trial. in this paper, we present data from two novel behavioral experiments that we designed to investigate the influence of quds on participants’ interpretation of conditionals as biconditionals. more precisely, we aim to test whether a qud that puts the focus of the conversation on the antecedent of the conditional, by asking about the conditions for the consequent (qud: will-q), positively influences a biconditional interpretation of the conditional, as compared to a qud that puts the focus of the conversation on the consequent by asking about the consequences of the antecedent (qud: if-p). in both experiments, participants are shown scenes of toy blocks together with a dialog between two characters that consists of a question, the qud, and a conditional answer. participants’ task is to select the scene(s) that they believe to be best described by the conditional. the set of scenes among which participants have to choose contains at least two scenes, an exhaustive, and a non-exhaustive situation, as we call them. both situations respectively contain (possibly among others) a blue and a green block, one in the upper left, the other in the lower right of the scene, where the falling of the upper block causes the lower block to fall as well; see figure 1 for the critical situations from experiment 2. since the conditional answer is always the same — “if the upper left block falls, the lower block will fall” where‘upper left’ and ‘lower’ are replaced by the respective color (green or blue) — we refer to the upper left block as the ‘antblock’, mentioned in the antecedent, and to the lower block as the ‘cons-block’, mentioned in the consequent. the two situations are manipulated with respect to the number of conceivable causes for the cons-block to fall. while, in both situations, the cons-block will fall if the ant-block falls, only in the non-exhaustive situation there is a second possible reason for the cons-block to fall: either because of its own position on the edge of the platform (condition internal, figure 1 stimulus a/c) or because of the falling of an additional block (condition external, figure 1 stimulus b/d). the idea is that when participants interpret the conditional “if the ant-block falls, the cons-block proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 106 https://doi.org/10.3765/elm https://www.elm-conference.net/ will fall” as biconditional, they should prefer to select the exhaustive situation as better described by the conditional. the difference between the two experiments concerns the concrete setup as we will explain below. previous studies have tested a qud-effect on the occurrence of cp as proposed by von fintel (2001), though yielding conflicting results (cariani & rips 2016, farr 2011, cariani & rips forthcoming). farr’s (2011) results provide quite strong evidence for the hypothesized effect of the qud, whereas only a minute effect, if any at all, was found by cariani & rips (2016). in the experiment from farr, participants read short vignettes and were asked whether a given conditional (e.g., “if you sell an eel, you get 2.50 euros”) is a sufficient answer to a question, encoding the qud (e.g., “what happens if i sell an eel?” vs.“ when do i get 2.50 euros?”). it may be considered problematic to ask for the sufficiency of the conditional answer (e.g., see cariani & rips 2016, lópez astorga 2014): the vignettes describe two possibilities to achieve the consequent (e.g., both, an eel and a pike cost 2.50), so that a no-answer to the question “did sahra answer kerstin’s question sufficiently?” does not necessarily imply — even though it may strongly suggest — that participants interpreted sahra’s conditional answer (“if you sell an eel, you get 2.50 euros”) to kerstin’s question (“when do i get 2.50 euros?”, qud: will-q) as biconditional. cariani & rips (2016) avoid this problem by measuring participants’ endorsement rates of the four inferences mentioned above (mp, mt, ac, da) to investigate the degree to which participants’ interpreted a conditional as biconditional. however, the experimental stimuli from cariani & rips (2016) — again short vignettes — come along with world knowledge that is hard to control for. for instance, in one of their trials participants learned that ‘john has taken a test on chapters 4-6 that has not been graded yet’. they were then asked whether the conditional (that they were told was true) ‘if john understood chapter 5, then john did well on the test’ implies that ‘john understood chapter 5’ when they also know that ‘john did well on the test’ (testing ac). as cariani & rips note themselves, participants might assume that the conditional simply does not state all conceivable conditions for the consequent; in this example, it is hard not to think of other reasons why john could do well on the test without having understood chapter 5 (e.g. by cheating), which may have influenced participants’ responses. an advantage of our non-text-based stimuli is that they allow to control participants’ elicited beliefs about the situations at hand much better. by showing participants animations of how the blocks behave, we hope to reduce the influence of additional world knowledge further. also, our measurement for how the conditional is interpreted does not involve a direct question about the sufficiency of the conditional as answer to the qud; we make participants select the situation in which they believe the conditional to be more appropriate. to anticipate our results, we find some evidence for an influence of the qud in the predicted way, yet the qud cannot fully explain the observed data. the results are nonetheless interesting as they suggest other aspects to play a role for the occurrence of cp, like the nature of the potential alternative causes (external vs. internal), leading to different sets of alternative utterances as well as different possible types of interpretations (causal vs. epistemic). 2. experiment 1. we preregistered the experiment based on a pilot study, in which we collected and analyzed data from 25 subjects. the code and the preregistration report are available on osf.3 3https://osf.io/47w85, https://tinyurl.com/255yaztv proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 107 https://doi.org/10.3765/elm https://www.elm-conference.net/ participants a total of 300 participants were recruited via the online platform prolific, including the 25 participants from our pilot study.4 all of them were self-reported native english speakers, at least 18 years old and had an approval rate on prolific of at least 80%. the cleaned data (see below) comprises 282 participants (103 male, 175 female, 3 other, 1 not specified) with a mean age of 32.8 years (range 18 – 84). for their participation, each participant received £1.25. setup & materials the experiment consisted of a training phase with 7 trials and a testing phase with 18 trials. 12 of the test trials were critical trials, the remaining 6 were control trials including 3 attention-checks. in the training phase, participants saw animations of block arrangements that were created with the rigid body physics engine‘matter.js’.5 the pictures shown in the test phase are screenshots (800× 500 pixel) of the corresponding animations right before they would start. manipulations. the manipulated variables comprise the qud as encoded in ann’s question and the shown pair of situations (exhaustive/non-exhaustive). the qud has the following three levels: neutral: “which blocks do you think will fall?”, if-p: “what happens if the ant-block falls?” and will-q: “will the cons-block fall?”. the exhaustive and the non-exhaustive situation have two levels each: the former either contains or does not contain a yellow distractor block in the upper right (exhaustive: with distractor, w/o distractor) and in the latter, the second cause why the consblock might fall is either due to its own position (non-exhaustive: internal) or due to the falling of a third block (non-exhaustive: external), as shown in figure 1.6 training phase the main purpose of the training phase was to familiarize participants with the physical properties of the blocks. to induce a maximal degree of uncertainty in the critical test trials about whether the ant-block will fall, the blocks in the training trials are all positioned such that it should be quite easy to judge whether they will fall, in particular after having seen a few examples.7 contrary to that, in the critical trials of the test phase the center of the ant-block lies exactly on the edge of the platform. the order of the training trials was randomized within participants, but always alternated between trials where some blocks fall and trials where nothing happens. in each training trial participants were first asked to select all blocks that they believed to fall, by clicking on buttons with the respective block icons (or saying ‘none’). only then, they were able to click on ‘run’ to start the animation to see which blocks actually fall and whether their selection was correct. we explicitly asked participants to look at all blocks shown in the scene to encourage them to consider the potential influence of each block on any other block. test phase in the test phase participants read a dialog between two characters, consisting of a question from ann and an answer to that question from bob. after reading the dialog, participants were shown two pictures of block arrangements and were asked to select the one that they rated as 4note that, since two prolific ids were erroneously recorded twice, we eventually recorded 302 instead of the initially planned 300 participants such that all 300 data sets are ensured to come from 300 distinct participants. from the two data sets that were associated with the same prolific id, the one with the later time stamp was excluded. 5https://brm.io/matter-js/ 6in the exhaustive situations, the distractor block never moves and has no influence on the falling of the other blocks. in the non-exhaustive situations, participants learn in the training phase that the cons-block falls if the antecedentor the yellow block falls or if both fall. 780% of the participants responded correctly in their respectively last trial of the training phase, whereas only about 50% gave the correct answer in the first training trial. proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 108 https://doi.org/10.3765/elm https://www.elm-conference.net/ more likely described by bob. 2.1. results. data exclusion we excluded all data from participants who did not give the correct response in: (i) all three attention-check trials or in (ii) more than one of the control trials or in (iii) the example test trial in the end of the training phase. additionally, we excluded the data from two participants whose comments in the end of the experiment indicated that they did not do the experiment properly. overall, the data of 282 participants remained to be included in the analysis. behavioral data figure 2 shows the proportion of participants who selected the exhaustive situation as the situation that is more likely described by bob’s conditional answer to ann’s question. participants’ choices seem to depend strongly on the stimulus (color coded); while similar results are observed for stimuli b and d on the one hand and stimuli a and c on the other hand, there is a striking difference between the results for stimuli b and d as compared to the results for a and c. the difference between stimuli a/c and b/d lies in the second cause for the cons-block to fall in the non-exhaustive situation: for stimulus b and d, it is the potential falling of an additional block whereas for stimulus a and c, it is the position of the cons-block itself that may cause it to fall. in the former two stimuli, participants show a strong preference for the exhaustive situation and in the latter two stimuli, participants seem to prefer the non-exhaustive situation (selection rates consistently below 0.5). the difference in responses observed between quds is much less striking. by eyeballing the data, we observe the expected tendency for stimulus a and c towards higher selection rates of the exhaustive situation when qud=will-q as compared to qud=if-p; for stimulus b and d, the same tendency is observed, but much less pronounced. statistical model we run a bayesian logistic regression model, using the r-package brms (bürkner 2017), with the qud, the stimulus (pair of two situations) and their interaction as predictors. as random effects, we include varying intercepts and slopes per participant and use brms default priors for all parameters. only for stimulus a, there is good reason to belief that the selection rate of the exhaustive situation will be larger when qud=will-q as compared to qud=if-p, reaching a posterior probability of approximately 96% (p (βqudif-p + βstimulusc:qudif-p < βqudwill-q + βstimulusc:qudwill-q) = 0.959, 95% ci: [-1.08, -0.03]). for the remaining three stimuli the posterior probabilities and 95% credible intervals for the respective comparison of parameters are 0.71 ([-1.01, 0.48], stimulus b), 0.833 ([-0.89, 0.23], stimulus c) and 0.712 ([-0.80, 0.38], stimulus d). the estimated posterior probability for the selection rate of the exhaustive picture to be larger when qud = will-q compared to when qud = if-p across all four stimuli amounts to 0.953. 2.2. discussion. we found supporting evidence for the postulated effect of the qud only for stimulus a, where the posterior probability for the exhaustive situation being selected more often when qud=will-q as compared to qud=if-p is reasonable large. however, the results for stimulus b,c and d also show a tendency towards this effect. these results may be related to the unexpectedly strong difference in participants’ responses between the four stimuli. especially salient is the systematic preference for the exhaustive situation in stimuli b and d (non-exhaustive: external) and for the non-exhaustive situation in stimuli a and c (non-exhaustive: internal). put differently, the biconditional interpretation is overall not proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 109 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 2: relative frequency of the selection of the exhaustive situation in experiment 1 as better described by the conditional “if the ant-block falls, the cons-block will fall”, for each qud, color coded by stimulus. errorbars are 95% bootstrap confidence intervals. very prominent for stimulus a and c, but it is for stimulus b and d. one aspect that might have influenced this difference is the set of salient alternative utterances: for stimulus b and d (external cause), there is a salient alternative to describe the dispreferred non-exhaustive situation, namely “if the yellow or the ant-block falls, the cons-block will fall”, which may help explain the selection rates of the exhaustive situation in stimulus b and d close to ceiling. for stimulus a and c, for which the selection rate of the exhaustive situation is much lower throughout all quds, there is no similarly salient alternative for the non-exhaustive situation where the cons-block may fall without the influence of any other block. further, for stimulus b and c, the observed preferences (exhaustive for b, non-exhaustive for c) may be strengthened by the presence of the yellow distractor block in only one of the two shown situations; participants may have favored the situation without the yellow block — corresponding to the respectively preferred situations — just because it is not mentioned in the conditional. this is not per se problematic to test for an effect of the qud, at least as long as the potential effect is not superposed completely by selection rates close to ceiling which we do observe for stimulus b. another possibility that may have influenced the results, in particular the preference for the non-exhaustive situation in stimuli a and c, is an epistemic instead of a purely causal interpretation of the conditional in these cases.8 assuming an epistemic interpretation, the conditional is particularly true in the non-exhaustive situation — which is overall preferred in stimuli a and c. in order to circumvent the possibility that a putative effect is not found due to participants’ 8epistemic interpretation in the sense that if the “nature of” the ant-block is such that it falls, the cons-block should fall as well since it is of the same “nature” as the ant-block: both blocks are identically centered on the edge of their platforms. proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 110 https://doi.org/10.3765/elm https://www.elm-conference.net/ strong tendency to prefer either the exhaustive or the non-exhaustive situation depending on the concrete scenes, we conducted a follow-up experiment where we do not force participants to choose one among two situations but allow them to select both. when qud=if-p, the conditional p → q, should in fact be accepted as description for both situations since other possible reasons for the consequent are simply expected to be irrelevant and thus, the conditional is an appropriate answer in the exhaustive as well as in the non-exhaustive situation. indeed, several participants in experiment 1 mentioned in their comments in the end of the study that for some trials, both pictures were possible. 3. experiment 2. the code and the preregistration report for experiment 2 are available on osf.9 participants 315 participants were recruited via the online platform prolific, using the same eligibility criteria as for experiment 1.10 the cleaned data (see below) comprises data from 181 participants (76 male, 103 female, 2 other) with a mean age of 37.12 years (range 18 – 68). for their participation, each participant received £1.67; on average they finished the experiment in approximately 14 minutes (range 4.5 – 46). setup & materials the training phase consisted of 8 trials and was followed by the test phase consisting of 17 trials split into 3 blocks, a practice block with 4 trials followed by two test blocks with 7 and 6 trials respectively. the trials of the two test blocks alternated between filler and critical trials and included an attention check trial after the first test block. the order of trials within filler, critical and practice trials was randomized for each participant. after each block, participants had the possibility to take a break before proceeding with the next block. in the end, we further asked participants to answer a set of questions about the experiment to (i) verify that they did not ignore ann’s question and (ii) to get an idea of how certain participants had to be such that they would select only one scene. procedure the most important difference in the procedure of experiment 2 compared to experiment 1 is that in experiment 2, participants were not forced to select a single picture.they read the same dialog as in experiment 1 but were shown three instead of two scenes to choose from. the third picture is referred to as the control scene since it shows a situation that contradicts bob’s conditional answer and should, thus, never be selected. by telling participants that ann sees part of the scene that bob describes, ann’s question was meant to be more purposeful: she seeks to get more information about a scene that she only has partial access to, see figure 3 for an example trial. while the partial scene was immediately visible in each trial, ann’s question, bob’s response and the three scenes had to be revealed one by one.with this setup we hoped to enforce participants to process both, the qud and the conditional, before making a selection. further, experiment 2 was built up as a game where participants can earn points, with the aim to incentivize that participants do not always select a single situation (or always both): when selecting two scenes, they get 50 points if the correct one is among them, otherwise, they loose 9https://tinyurl.com/pphck52h, https://tinyurl.com/pphck52h 10we increased the number of recorded participants by 100 after having cleaned the data of the originally planned 215 participants since we had to exclude many more participants than expected who did not fulfill our predefined criteria. all data was analyzed only after all 315 participants were recorded. proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 111 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 3: critical test trial from experiment 2 where qud=if-p, non-exhaustive (right) external and exhaustive (left) w/o distractor. the picture in the middle is the control scene. 100 points. selecting only the correct scene is awarded with 100 points, but when a single scene is selected that is not the correct one, participants lose 100 points. therefore, in the long run, participants are better off to select two scenes when they are undecided. manipulations as in experiment 1, we manipulate the qud encoded in ann’s question, but without using the neutral condition (“which blocks do you think will fall?”) in the critical test trials. while the four critical stimuli (pairs of an exhaustive and a non-exhaustive situation) are the same as in experiment 1, there are only 6 critical trials in experiment 2: since ann’s question, the qud, relates to the block that is shown in the partial scene (when qud=if-p, the partial scene shows the ant-block, when qud=will-q, it shows the cons-block), it should be at the same position as the respective block in the three pictures of scenes among which participants make their selection; only then none of the three situations can be excluded just because it does not match the part of the scene that ann sees. thus, stimuli a / c are not combined with qud will-q. training phase the animations in the training phase were the same as in experiment 1 plus one additional trial which showed the situation that contradicts the conditional “if the ant-block falls, the cons-block will fall” (see figure 3, middle), which is the control scene in the critical conditions of the test phase. test phase the test phase consists of three blocks, a practice and two test blocks. the practice block consists of 4 trials in which participants got feedback about the correct picture and the number of points they received with their selection.the main purpose of the practice trials was to demonstrate that bob’s responses are informative and to make participants learn how their choices impact the amount of points they get. the procedure in the the trials of the two test blocks was the proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 112 https://doi.org/10.3765/elm https://www.elm-conference.net/ c d a b non−exhaustive both exhaustive non−exhaustive both exhaustive non−exhaustive both exhaustive non−exhaustive both exhaustive 0.0 0.2 0.4 0.6 0.0 0.2 0.4 0.6 0.2 0.3 0.4 0.5 0.6 0.2 0.4 0.6 qud se le ct io n ra te qud if−p will−q figure 4: relative frequency of the scene(s) participants selected in experiment 2 based on the conditional “if the ant-block falls, the cons-block will fall”; color code represents the two quds, errorbars are 95% bootstrapped confidence intervals. same as in the practice block, except that participants did not receive feedback anymore. in order to keep the character of the game up without influencing participants’ choices, they were told that they would get their final score in the end of the experiment. the filler trials in the test blocks were designed such that there is a similar number of trials where qud=if-p (6) and qud=will-q (5). 3.1. results. data exclusion we excluded all data from participants who fulfilled at least one of the following criteria: (i) they did not select the correct scene in the attention-check trial, (ii) they selected the control scene at least once in the test phase (excluding the practice block), (iii) they affirmed either that they only read bob’s answer, but not ann’s question, or that ann’s question was always the same or (iv) they responded within less than 6 seconds in at least 2 of the critical trials. behavioral data figure 4 shows participants’ average responses in experiment 2 for each of the four critical stimuli. similarly to experiment 1, we observe a preference away from an exhaustive interpretation in stimuli a and c, where selecting both situations is much more likely than selecting only one situation; in stimulus c, selecting only the non-exhaustive situation is even more likely than selecting only the exhaustive situation. contrary to that, stimuli b and d again show a preference towards an exhaustive interpretation: the selection rate of only the exhaustive situation is much higher in these stimuli than it is in stimuli a and c. selecting both situations is, however, almost equally likely for stimuli b and d compared to a and c. by eyeballing the results in figure 4, the qud will-q shows the predicted effect for stimulus b: we observe an increase in the selection of the exhaustive situation when qud=will-q compared to qud=if-p and a decrease in the selection of both situations. for stimulus d, selecting both proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 113 https://doi.org/10.3765/elm https://www.elm-conference.net/ situations is more likely than selecting only the exhaustive situation when qud=if-p, but we do not observe the predicted effect of the qud on the selection rate of the exhaustive situation, which is in fact on average more likely when qud = will-q but the increase seems to be marginal. statistical model we run an ordinal regression model with brms (bürkner & vuorre 2019, bürkner 2017) with the qud, the non-exhaustive and exhaustive situation and the interaction between the qud and the exhaustive situation and between the exhaustive and the non-exhaustive situation as predictors, including by-participants random intercepts and slopes for each predictor except for interactions. we chose an ordinal regression model as the three response categories reflect the degree of how exhaustive the conditional is interpreted: the selection of only the nonexhaustive situation corresponds to a maximally non-exhaustive interpretation and the selection of only the exhaustive situation to a maximally exhaustive interpretation, selecting both situations corresponds to an interpretation in between both extremes. for stimulus b, the posterior probability that participants interpret the conditional more exhaustively when qud=will-q as compared to qud=if-p amounts to 93%(p (βqudwill-q + βexhwod + βqudwill-q:exhwod > βexhwod) = 0.93, 95% ci: [-0.04, 0.58]), pointing towards the postulated effect of the qud. as figure 4 suggested, for stimulus d, the posterior probability is much lower (p (βqudwill-q > 0) = 0.56, 95% ci: [-0.26, 0.32]) and does not provide evidence for our hypothesis. we further speculated (a) that when qud=will-q, the selection rate of only the exhaustive situation will be larger than the selection rate of both situations and (b) when qud=if-p, the selection rate of both situations will be larger than of only the exhaustive situation. our data only provides evidence for (b) in stimuli a and c (posterior probability for p (both | qud=if-p) > p (exhaustive | qud=if-p) is 1). 3.2. discussion. only the data for stimulus b provides a reason to believe in a more exhaustive interpretation of the conditional when qud=will-q as compard to qud=if-p, for stimulus d, the qud does not seem to have the same alleged effect. when we look at the exact conditions in which participants responses differ between stimuli b and d, we find the strongest difference in the selection rate of the exhaustive situation for stimuli b and d when qud=will-q, with a posterior probability for p (e | qud = will-q, stimulus = b) > p (e | qud = will-q, stimulus = d) of 0.92. when qud=if-p, the posterior probability that p (e | qud = if-p, stimulus = b) > p (e | qud = if-p, stimulus = d) is 0.52. assuming that potential alternative causes are irrelevant for the choice that participants make when qud = if-p, it seems reasonable that the observed difference between participants’ responses in stimuli b and d can mainly be ascribed to the condition where qud=will-q since when qud=if-p, the focus of the conversation lies on the consequences of the antecedent and so, participants are not expected to consider other potential causes for the consequent. further, it may have been the case that the presence of the distractor block in the exhaustive situation of stimulus d (that is absent in b) made participants more hesitant to decide for a single scene and thus they tend to choose both situations more often in this stimulus when alternative causes are considered, i.e., when qud=will-q. participants were indeed encouraged to select a single situation only when they were very confident, which 86% of participants confirmed in the questions in the end of the experiment. concerning the results for stimuli a and c, we observe a tendency towards a non-exhaustive proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 114 https://doi.org/10.3765/elm https://www.elm-conference.net/ interpretation of the conditional (selection of the non-exhaustive situation or both), similarly to what we saw in experiment 1. in fact, for stimulus c, the posterior probability of the probability to select only the the non-exhaustive situation to be larger than the probability to select only the exhaustive situation amounts to almost 0.98 (p (non-exhaustive | qud=if-p, stimulus = c) > p (exhaustive | qud=if-p, stimulus = c)). the fact that in experiment 2 we find that if participants choose a single scene in stimulus c, they select the non-exhaustive situation significantly more often than the exhaustive situation — which is not the case for stimulus a — might help to explain why, in experiment 1, we did not find the predicted qud-effect for stimulus c which we found for stimulus a; a general tendency towards the non-exhaustive situation may have interfered with a putative qud-effect which may thereby become harder to find, especially when we assume that this preference would only become stronger when the alternative causes are assumed to be particularly considered, that is, when qud=will-q. 4. conclusion. overall, we find some evidence supporting the hypothesis that a qud that focuses on the conditions bringing about the consequent yields a more exhaustive interpretation of an indicative conditional than does a qud that focuses on the consequences of the antecedent. our results are far from being conclusive, yet they show that the interpretation of conditionals as biconditionals is likely to be the result of an interplay of various factors. especially experiment 1 showed that when forced to choose either an exhaustive or a nonexhaustive situation, participants showed substantially different preferences depending on the nature of the second conceivable cause for the cons-block to fall in the non-exhaustive situation. it either fell because of a third block (stimuli b/d) or because of its own position on the edge of a platform (stimuli a/c). stimuli b/d yield a strong preference of the exhaustive-situation whereas in stimuli a/c we observed a (less strong) preference of the non-exhaustive situation — across all quds including a neutral question. the second cause for the cons-block to fall as it is is realized in stimuli a and c, namely because of its own position on the edge, comes along with another possible interpretation of the conditional that does not apply to stimuli b and d: the conditional may receive an epistemic instead of a causal interpretation. consider the following conditional as an example for a conditional describing a situation similar to those in stimuli a and c, receiving an epistemic interpretation: “if that guy solved the puzzle, she will solve it [too / all the more]”. it does not seem to suggest that ‘if that guy does not solve the puzzle, she will not solve it’, rather, it suggests that ‘she might solve it, while he might not, but if he does, she will as well’. and this seems to be the case even if the conditional is an answer to the question “will she solve the puzzle?”. in other words, under an epistemic interpretation of the conditional we should not expect to see a difference between the quds whereas we do expect a difference when the conditional receives a non-epistemic, causal interpretation. therefore, disentangling a causal versus an epistemic interpretation of the conditional may help to get a cleaner picture of what is going on here. further, the observed tight connection between conditionals and causality generally suggests that it may be worth to look at the production of conditionals in comparison to the use of causal language (e.g., “x may make y fall” or “y may fall because of x”) to learn more about how participants use (and interpret) conditionals. proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 115 https://doi.org/10.3765/elm https://www.elm-conference.net/ references atlas, jay david & stephen c. levinson. 1981. it-clefts, informativeness and logical form: radical pragmatics (revised standard version). in radical pragmatics, 1–62. academic press. bürkner, paul-christian. 2017. brms: an r package for bayesian multilevel models using stan. journal of statistical software 80(1). 10.18637/jss.v080.i01. bürkner, paul-christian & matti vuorre. 2019. ordinal regression models in psychology: a tutorial. advances in methods and practices in psychological science 2(1). 77–101. 10.1177/2515245918823199. cariani, fabrizio & lance j rips. 2016. experimenting with (conditional) perfection: tests of the exhaustivity theory . cariani, fabrizio & lance j rips. forthcoming. experimenting with (conditional) perfection: tests of the exhaustivity theory. in stefan kaufmann, over david & ghanshyam sharma (eds.), conditionals: logic, linguistics, and psychology, palgrave. cummins, denise d., todd lubart, olaf alksnis & robert rist. 1991. conditional reasoning and causation. memory & cognition 19(3). 274–282. 10.3758/bf03211151. de cornulier, benoit. 1983. ‘if’and the presumption of exhaustivity. journal of pragmatics 7(3). 247–249. ducrot, oswald. 1969. présupposés et sous-entendus. langue française 10.3406/lfr.1969.5456. evans, jonathan st bc, jonathan st bt evans, stephen e. newstead & ruth mj byrne. 1993. human reasoning: the psychology of deduction. psychology press. farr, marie-christine. 2011. focus influences the presence of conditional perfection: experimental evidence. proceedings of sinn & bedeutung 15. 225–239. fillenbaum, samuel. 1986. the use of conditionals in inducements and deterrents. in alice ter meulen, charles a. ferguson, elizabeth closs traugott & judy snitzer reilly (eds.), on conditionals, 179–196. cambridge: cambridge university press. 10.1017/cbo9780511753466.010. geis, michael l. & arnold m. zwicky. 1971. on invited inferences. linguistic inquiry 2(4). 561–566. horn, laurence r. 2000. from if to iff: conditional perfection as pragmatic strengthening. journal of pragmatics 32(2ooo). 289–326. lilje, gerald w. 1972. uninvited inferences. linguistic inquiry 540–542. lópez astorga, miguel. 2014. ¿podemos evitar la perfección del condicional enfocando el antecedente o son necesarios antecedentes alternativos? revista signos 47(85). 267–292. markovits, henry. 1986. familiarity effects in conditional reasoning. journal of educational psychology 78(6). 492. van canegem-ardijns, ingrid & william van belle. 2008. conditionals and types of conditional perfection. journal of pragmatics 40(2). 349–376. 10.1016/j.pragma.2006.11.007. van der auwera, johan. 1997. pragmatics in the last quarter century: the case of conditional perfection. journal of pragmatics 27(3). 261–274. von fintel, kai. 2001. conditional strengthening. unpublished manuscript . proceedings of elm 2: 104-116, 2023 britta grusdt, mingya liu and michael franke: testing the influence of quds on the occurance of conditional perfection. 116 https://doi.org/10.3765/elm https://www.elm-conference.net/ nonboolean conditionals paolo santorio & alexis wellwood∗ abstract. on standard analyses, indicative conditionals (ics) behave in a boolean fashion when interacting with and and or. we test this prediction by investigating probability judgments about sentences of the form ⌜a → b { and/or } c → d⌝. our findings are incompatible with a boolean picture. this is challenging for standard analyses of ics, as well as for several nonclassical analyses. some trivalent theories, conversely, may account for the data. keywords. indicative conditionals, connectives, probability, likelihood estimation, trivalence 1. introduction. our topic is the interaction between indicative conditionals, i.e. sentences like (1), and sentential connectives like and and or. (1) if alia flipped the coin, the coin landed heads. on classical accounts of indicative conditionals (henceforth, ics), such as stalnaker’s (1968, 1970, 1975) and kratzer’s (1981, 1986, 2012), ics are analyzed as modalized claims with epistemic flavor. for example, kratzer-style truth conditions for (1) are in (2). (2) ⟦(1)⟧w = true iff, for every world w′ that is compatible with the speaker’s evidence in w and such that alia flipped the coin in w′, the coin landed heads in w′ these analyses, paired with standard analyses of and and or, predict that ics interact with the latter in a boolean way. the truth values of conjunctions and disjunctions involving ics, such as (3), are determined via the truth tables for the connectives ‘∧’ and ‘∨’. (3) if alia flipped the coin, the coin landed heads, { and/or }, if billy tossed the die, the die landed even. truth-conditional analyses of ics are controversial, and a number of alternative analyses have been proposed.1 we don’t have space to survey these accounts, but it is worth pointing out that, despite their nonclassical bent, most of them are still boolean. this paper aims to test experimentally the prediction that ics are boolean. to do this, we investigate probability judgments about sentences of the form ⌜a→ b { and/or } c→ ∗thanks to audiences at the semantics and pragmatics of conditional connectives workshop at the 43rd deutsche gesellschaft für sprachwissenschaft (dgfs) in freiburg, the university of maryland at college park, and elm 2 at upenn. thanks to two anonymous elm referees for very useful comments. authors: paolo santorio, university of maryland, college park (santorio@umd.edu) & alexis wellwood, university of southern california (wellwood@usc.edu). 1here are just a couple of nonclassical options. dynamic/informational analyses (see e.g. gillies 2004, 2009) treat ics as imposing constraints on information states rather than expressing propositions, with the goal of capturing nonclassical logical features. sequence-based analyses (van fraassen 1976, kaufmann 2009, khoo 2022, santorio 2022 to mention a few) assign ics semantic values that are more fine-grained than worlds, with the goal of vindicating bridge principles between ics and probability. proceedings of elm 2: 252-264, 2023 c©2023 paolo santorio and alexis wellwood published by the lsa with permission of the author(s) under a cc by license. 252 https://doi.org/10.3765/elm https://www.elm-conference.net/ d⌝. as we will show, our findings suggest that, in some cases, the interaction between ics and connectives is nonboolean. while there is much room for further investigation, such results challenge both standard truth-conditional theories and their non-truthconditional cousins that still uphold that ics are boolean. in the final section, we briefly suggest that trivalent theories of ics are well-placed to predict the data. we proceed as follows: §2 introduces relevant theoretical background, §3 describes our experiments, and §4 sketches a trivalent analysis of conditionals that is able to capture the data. 2. background: connectives and probability. on classical theories, the meaning of the sentential connectives and and or in natural language is captured by the truth tables of the corresponding boolean connectives ‘∧’ and ‘∨’ in first-order logic (see table 1). a b a ∧b t t t t f f f t f f f f a b a ∨b t t t t f t f t t f f f table 1: classical truth tables for ‘∧’ and ‘∨’. boolean interpretations of and and or entail constraints about probability. consider and. since a∧b is true when and only when both conjuncts are true, the probability of a conjunction is a lower bound on the probability of each conjunct: the probability of a is at least as high as the probability of a∧ b, and possibly higher. moreover, if a does not entail b, the probability of a is strictly greater than the probability of a∧ b, since there are some a-possibilities that are not b-possibilities, and hence not a∧b-possibilities (see figure 1, left). analogous facts hold for or, mutatis mutandis. a a∧b a∨b a figure 1: diagrams illustrating the validity of and-drop and or-drop. in sum, the following principles hold for all natural language sentences that express propositions, on the assumption that connectives are boolean:2 and-drop. if a ⊭ b, p r(a) > p r(a∧b) or-drop. if a ⊭ b, p r(a∨b) > p r(a) 2the notion of entailment used in the principles can be interpreted as purely logical entailment, or as contextual entailment (i.e. entailment given what is known in the context). in the former case, the notion of probability that is relevant for the two principles is a logical notion of probability, in the latter it is a notion of credence. (for interpretations of probability, see hájek 2019.) proceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 253 https://doi.org/10.3765/elm https://www.elm-conference.net/ for an illustration, suppose that a fair six-sided die is tossed, and consider (4) and (5). (4) if the die landed odd, it landed on 1. (5) if the die landed odd, it landed on 1, and, if it landed even, it landed on 2. given and-drop, and given that (very plausibly) the two ics in (5) do not entail each other, classical analyses predict the probability of (4) should be higher than that of (5). the intuition driving the current paper is that ics trigger failures of and-drop and or-drop. just (4) and (5) provide an illustration. consider (4) first. since there are three equiprobable cases in which the die lands on an odd number, the intuitive probability of (4) is 1/3. via and-drop, we would expect the probability of (5) to be lower. yet, intuitively, the probability of (5) is also 1/3. to see this, notice that (5) is intuitively equivalent to (6). (6) the die landed on 1 or 2. since (6) is true in two of six equiprobable cases, its probability is also 1/3. but then, given the equivalence between (5) and (6), the probability of (5) should also be 1/3. so there is at least a first pass intuitive case for the invalidity of and-drop. but we don’t think that the empirical situation can be settled by these simple judgments. first, and-drop is a basic principle, and it takes very solid data to reject it. second, the case that we just discussed appears to rely on controversial assumptions about the probabilities of ics. in particular, we assumed that the probability of (4) equals the relevant conditional probability, in conformity to so-called stalnaker’s thesis. stalnaker’s thesis. for any a and b such that p r(a) > 0: p r(if a, b) = p r(b | a) stalnaker’s thesis is highly intuitive, as is the probability judgment about (4). but stalnaker’s thesis is also notoriously problematic, giving rise to so-called triviality results in combination with fairly minimal assumptions.3 we intend to sidestep any assumptions about stalnaker’s thesis and probabilities of conditionals. we will show that, even without these assumptions, we can find evidence for the invalidity of and-drop and or-drop. 3. experiments. we set out to test the widely-held view that conjunctions and disjunctions of ics are boolean. if this view is correct, conjunctions of ics should be estimated at a lower probability than either of their conjuncts, and disjunctions of ics should be estimated at a higher probability than either of their conjuncts. we chose our experimental methodology in response to two major desiderata. (i) we wanted it to be as simple as possible, given the complexity of the sentences to be tested. (ii) we wanted to avoid making assumptions about the probabilities of the relevant uncoordinated conditionals, so that we didn’t need to rely on stalnaker’s thesis. in our 3the first triviality result was presented by david lewis in his 1976. for other triviality results, see bradley 2000, bradley 2007, charlow 2016, russell & hawthorne 2016. see hájek & hall 1994 and khoo & santorio 2018 for a classical and a more recent overview. proceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 254 https://doi.org/10.3765/elm https://www.elm-conference.net/ task, participants were presented with simple shapes whose color changes to one of two colors in dynamically-unfolding scenes. we then let them generate their own likelihood estimations for the conditionals, based on the frequencies of the observed events. so our experiment makes no normative assumptions about what the ‘correct’ likelihoods of certain statements are supposed to be. 3.1. experiment 1. we ran a likelihood estimation task featuring different frequencies of simple event types and asked whether participants’ responses to simple and complex ics about those event types showed boolean behavior. 3.1.1. participants. we recruited 200 participants on amazon’s mechanical turk (mturk) platform in accordance with a protocol approved by usc’s institutional review board. participation selection was restricted to individuals with ip addresses located in the united states, whose hit approval rate was greater than or equal to 99%, and whose number of approved hits was greater or equal to 1000. we did not require master status. our experiment design included an attention check that allowed us to filter participants according to whether they passed or failed this check. based on this, we excluded 47 participants (23.5%) prior to data analysis, for a sample of 153 participants in experiment 1. 3.1.2. materials. we designed a set of 8 base animations involving two shapes–a square and a circle–turning one of four colors. the square always turned either red or yellow, and the circle always turned either green or blue. the depicted events involved one or two shapes “traveling” in a “car” into a tunnel; once the car entered the tunnel, the shape(s) changed one of the two colors (see the sample in figure 2, left). these base animations could felicitously be described as, e.g., the square turned red. we also designed a “mystery” animation in which the identity of the shape(s) were hidden (see figure 2, right). figure 2: event animations (left) and mystery animation (right) of experiment 1. 3.1.3. design. our task tested coordinated and uncoordinated conditional sentences in three phases. in a first experience phase, participants were exposed to a series of 24 events in random order, and asked for two binary judgments about sentences describing what happened in this phase (attention check; both intended to be judge true), (7)-(8).4 4if participants failed to answer “true” to both of these questions, this indicated to us that they failed to store the range of color changes for each of the shapes. these participants were subsequently excluded proceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 255 https://doi.org/10.3765/elm https://www.elm-conference.net/ in the second uncertainty phase, they were exposed to a “mystery” event and asked to judge the likelihood that a certain (uncoordinated) conditional sentence was true. finally, in the test phase, they were exposed to further “mystery” events and asked the likelihood of conjoined or disjoined conditionals. our design’s three factors–connective type, compatibility, and frequency–were manipulated across these phases. (7) if the square enters the tunnel, it always turns red or yellow. (8) if the circle enters the tunnel, it always turns green or blue. the first factor, connective type, was tested within subjects and concerned whether the target sentences involved a connective, and if so, which. the no connective condition was tested in the uncertainty phase: participants were exposed to 4 “mystery” animations, and asked to judge how likely each type of event was (square-red, square-yellow, circlegreen, circle-blue) using the sentences schematized in (9)-(10). the and and or conditions were tested in the test phase: participants were exposed to a second round of 4 mystery animations, and evaluated the coordinated ics schematized in (11)-(12). (9) if the car was carrying the square, the square turned { red, yellow }. (10) if the car was carrying the circle, the circle turned { green, blue }. (11) if the car was carrying the square, the square turned red { and / or } if the car was carrying the circle, the circle turned green. (12) if the car was carrying the square, the square turned yellow { and / or } if the car was carrying the circle, the circle turned blue. the second factor, frequency (50/50, 75/25), was tested between subjects and concerns the base rates in the experience phase of the types of events our sentences are about. in the 50/50 condition, the rate of the square turning red or yellow and of the circle turning green or blue were equal. in the 75/25 condition, the square turned red 75% of the time and the circle turned green 75% of the time. in the 50/50 condition, then, there was no distinction in frequency for (11) and (12). in the 75/25 condition, the targets in (11) counted as high frequency and those in (12) counted as low. the third factor, compatibility (compatible, incompatible), was also tested between subjects, and concerns whether the experience phase involved animations involving just one, or additionally two entities traveling in the car at a time. (the color changes of the two entities, and their relative frequencies, were independent of whether the entities traveled together or alone.) if the participants’ experiences showed that the car could carry two shapes, then the antecedents of the conditionals in our coordinated sentences (11)-(12) were compatible, i.e. they could both be true. if their experience showed that the car always carried only one shape, the antecedents were incompatible. 3.1.4. procedure. after accepting the hit on mturk, participants were instructed to click a link that would take them to the experiment hosted on google’s firebase platfrom analysis. proceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 256 https://doi.org/10.3765/elm https://www.elm-conference.net/ form. there, participants were welcomed to the study and presented with an instruction screens that explained the task and what they were being asked to do in overview. the experience phase was described as follows: “your task is to keep track of the frequency of the color changes for each shape. at the end of the series, you’ll be asked to judge simple statements about what you saw.” the message to keep track of the color changes and their frequency was reinforced on a second screen introducing this phase. following the randomly-ordered presentation of the 24 events, participants were presented with the attention check, evaluating (7)-(8) in random order in response to the question, “does the statement accurately describe what you know about the { square, circle }?” immediately following this, the uncertainty phase began, with participants first reminded that they would be seeing “mystery” animations, and that they would be asked to judge the likelihood of a statement about each one, “given what [they] learned in the first part of the study.” the statements were as in (9)-(10). there were 4 such trials. there were no additional instructions leading into the test phase; participants simply saw 4 more mystery animations and, after each, were asked about the coordinated conditionals schematized in (11)-(12). for all likelihood estimations in the uncertainty and test phases, participants were asked, “given the animation you just saw, how likely is it that the complex statement below is true?”, and had to click and drag a slider ranging from 0 (“completely unlikely”) to 100 (“completely likely”) to record their response. 3.1.5. results. we report the results of a 3x2x2 anova with a within-subject error term for connective type. we found that our participants overestimated input frequencies in the 50/50 condition (‘balanced’ inputs occurred 50% of the time, mean estimate 68%) and in the lower frequency events of the 75/25 condition (‘lower’ input 25%, estimate 46%; cp. ‘higher’ input 75%, estimate 75%). importantly for us, however, the ordering between estimates was accurate, and the 50/50 and 75/25 conditions were significantly different (f = 8.15, p < .005). probing this result further, we conducted pairwise t-tests with bonferroni adjustment on judgments between the 25%, 50%, and 75% inputs, and all were significantly different (ps < .001). crucially, however, participants’ likelihood estimates were not impacted by the factors connective type or compatibility (both ps > .53). the lack of effect of connective type (and compatibility) shows that uncoordinated ics were assigned, on average, the same probability as conjunctions and disjunctions of ics. yet we cannot simply attribute these results to task difficulty or the like, given the evidence that subjects tracked input frequencies broadly accurately. see figure 3. 3.1.6. discussion. the results of experiment 1 militate against the validity of and-drop and or-drop. these principles lead to the prediction that connective type would make a difference in likelihood estimations such that: the likelihoods assigned to uncoordinated ics should be higher than the likelihoods assigned to conjunctions of ics, and lower than the likelihoods assigned to disjunctions of ics. this is not what we found. if these results accurately reflect speakers’ knowledge of ics, then it turns out that ics aren’t necessarily boolean. we sketch in §4 how this may be accommodated. a potential worry with this experiment is that likelihood judgments can be unreliproceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 257 https://doi.org/10.3765/elm https://www.elm-conference.net/ figure 3: results of experiment 1. able in the best of cases, let alone in the case of complex coordinated ics. for example, some well-known results in the psychology of reasoning show that subjects can fall into ‘cognitive illusions’, which cause them to evaluate some conjunctions as more likely than conjuncts.5 one might worry, then, that the patterns we observed are due merely to distortions in judgments about likelihood, as opposed to issues in the semantics of ics. 3.2. experiment 2. this experiment tests whether our likelihood estimation task would generate boolean behavior under different linguistic circumstances. we presented participants with the same task as in experiment 1, with targeted modifications to support the evaluation of minimally-different, but non-conditional sentences. 3.2.1. participants. we recruited 100 participants on mturk with the same filter parameters as for experiment 1. our task incorporated the same binary attention check as did experiment 1, and based on the results of that check we excluded 17 people (17%), with 5the locus classicus for this claim is tversky & kahneman 1983. proceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 258 https://doi.org/10.3765/elm https://www.elm-conference.net/ the result that we report the data from 83 participants. 3.2.2. materials. we used the same set of 8 base animations as experiment 1, but our “mystery” animations were different. instead of obscuring the identity of the shape(s) in the car, those shapes were visible, and only their color change in the tunnel was obscured. 3.2.3. design. we modified the design of experiment 1 to accommodate testing nonconditional coordinated and disjoined sentences. here, the experience phase manipulated frequency as in experiment 1, and at the end of that phase, the attention check sentences (our basis for any exclusions from the participant pool) were as in (13)-(14). in experiment 2, though, we only used the distribution of animations corresponding to the compatible level of compatibility (i.e., a mixture of singleand double-entity animations). without the use of conditional sentences in this experiment, there was no manipulation that tested the joint (un)satisfaction of their antecedents. (13) the circle always turns blue or green. (14) the square always turns red or yellow. in the uncertainty phase, participants were presented with the modified mystery animations in which the identity of a single shape in the car was shown, but its color change was hidden. at the end of each mystery animation, they were asked to evaluate one of the 4 sentences schematized in (15)-(16). this corresponded to a test of the level ‘no connective’ level of connective type in this experiment. the label of the shape always matched the identity of the shape in the animation. (15) the square turned { red/yellow }. (16) the circle turned { green/blue }. finally, in the test phase, participants were presented with the modified mystery animations that showed the two shapes present in the car, but their color changes were hidden. following each of these mystery animations, they were asked about one of the 4 sentences schematized in (17), in random order. (17) the square turned { red/yellow } and the circle turned { green/blue }. 3.2.4. procedure. identical to experiment 1. 3.2.5. results. we report the results of a 3x2 anova with a within-subject error term for connective type, as in experiment 1. we did not observe a simple main effect of frequency in this experiment (p > .95), as the overall averages for the 50/50 and 75/25 conditions were quite close (50/50 64.7%, 75/25 64.9%; see figure 4). probing this further, we conducted pairwise t-tests with bonferroni correction to each level of input frequency, and found that each was significantly different from the others (all ps < .001), and, while there were distortions in the estimates, they were nonetheless appropriately ordered (input 25%, estimate 48.9%; input 50%, estimate 64.7%; input 75%, estimate 81.8%), as we previously observed for ics. here, however, we found a main effect of connective type (f = 11.7, p < .001). pairproceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 259 https://doi.org/10.3765/elm https://www.elm-conference.net/ wise t-tests with bonferroni correction between the levels of connective type revealed this result to reflect estimates for and differing significantly from or and no connective (and 58.4%, or 68.6%, none 66.1%; both ps < .007), while or and no connective didn’t differ (p = .34). this is suggestive of at least partially boolean behavior. we also found an interaction effect between connective type and frequency, f = 3.3, p = .04. unpacking this interaction, we found that the boolean pattern was stable in the 50/50 condition and minimized in the 75/25 condition. pairwise t-tests with bonferroni correction inside the subsets of the data corresponding to the 50/50 and 75/25 conditions revealed the following results. in the 50/50 condition, and was significantly different than or and no connective (both ps < .004), but or didn’t differ from no connective (p = .15). in the 75/25 condition, the connectives didn’t differ significantly from one another (all ps > .26). this shows expected boolean behavior at least in the 50/50 condition. figure 4: results of experiment 2. 3.2.6. discussion. in experiment 2, we observed the expected boolean behavior for nonics in the 50/50 condition. this alleviates, to some extent, concerns that our paradigm wouldn’t be sensitive enough to detect such behavior. of course, we recognize that these results don’t fully support the idea that, when conditionals are not involved, subjects give boolean judgments about our scenarios. more probing is needed, as we emphasize below. 4. general discussion. our experiment shows at least some evidence that and-drop and or-drop fail. we have pointed out that this finding is challenging for most theories of ics. but what theories can potentially accommodate it? as it turns out, some versions of so-called trivalent semantics of ics predict failures of and-drop and or-drop. the central idea of trivalent theories is simple, and goes back to de finetti (1936): ics have a truth value just in case their antecedent is true, and are undefined otherwise. this idea has been developed in a number of ways.6 here we 6for some recent views, see rothschild 2014, lassiter 2020. for an interesting and very detailed discussion of trivalent theories, see égré et al. 2021. proceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 260 https://doi.org/10.3765/elm https://www.elm-conference.net/ present a version of the trivalent theory that is inspired by bradley 2002. the key idea behind the semantics is that every clause has definedness conditions and truth conditions. we use ‘def(a)’ to denote the former and ‘true(a)’ to denote the latter. the semantic clauses for connectives are the following: ⟦¬a⟧ = { def. at w iff w ∈ def(a) true at w iff w < true(a) ⟦a→ b⟧ = { def. at w iff w ∈ true(a) and w ∈ def(b) true at w iff w ∈ true(a)∩true(b) ⟦a∧b⟧ =  def. at w iff w ∈ def(a) or w ∈ def(b) true at w iff: if w ∈ def(a), w ∈ true(a) and if w ∈ def(b), w ∈ true(b) ⟦a∨b⟧ =  def. at w iff w ∈ def(a) or w ∈ def(b) true at w iff: if w ∈ def(a) and w ∈ true(a) or w ∈ def(b) and w ∈ true(b) let us emphasize one point. the definedness condition for conjunction, which is borrowed from bradley 2002, is unusually weak: for a conjunction to be defined, all that is needed is that at least one of the conjuncts is defined. this is nonstandard, even for trivalent frameworks (see e.g. lassiter 2020 for a different definedness condition for). but it is crucial for predicting the failure of and-drop. since our language involves truth-value gaps, we have to adopt a non-bivalent notion of probability ptriv. to define the latter, we follow cantwell 2006 (see also lassiter 2020). the basic idea is that the trivalent probability ptriv of a sentence a can be defined from standard probabilities, in the following way: ptriv(a) equals the ratio of the probability of the truth of a, divided by the probability that a is defined.7 ptriv(a) = p r(true(a)) p r(def(a)) , if p r(def(a))> 0 we can easily show how, given this semantics and this way of defining ptriv, we can invalidate and-drop. consider the epistemic state of a subject who has just observed a ‘mystery’ animation in our experiment 1. suppose that the epistemic state of the subject includes four worlds. in w1 and w2 the car is carrying the circle, which turns green in w1 and blue in w2. in w3 and w4 the car is 7we assume, with cantwell 2006, that the standard kolmogorov axioms hold for the notion of probability p r. see also rothschild 2014 and lassiter 2020 for an alternative definitions of trivalent probability. proceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 261 https://doi.org/10.3765/elm https://www.elm-conference.net/ carrying the square, which turns yellow in w3 and red in w4. suppose moreover that these worlds are assigned equal probability by the subject, i.e. that we have p r(w1) = p r(w2) = p r(w3) = p r(w4) = 1/4. now consider the two ics: (18) if the car was carrying the circle, the circle turned blue. (19) if the car was carrying the square, the square turned red. (18) is defined at w1 and w2, and true at w2. so, by the definition of ptriv, we have: ptriv(18) = p r(true(18)) p r(def(18)) = p r({w2}) p r({w1,w2}) = 1/4 1/2 = 1/2 via analogous reasoning, we also have that ptriv(19) = 1/2. and now, consider the conjunction of (18) and (19): (20) if the car was carrying the circle, the circle turned blue, and if the car was carrying the square, the square turned red. (20) is defined at all worlds (since the left conjunct is defined at w1 and w2, the right conjunct is defined at w3 and w4, and conjunctions are defined at a world w iff at least one of the conjuncts is defined at w), and true at w2 and w4. so we have: ptriv(20) = p r(true(20)) p r(def(20)) = p r({w2,w4}) p r({w1,w2,w3,w4}) = 1/2 1 = 1/2 so we have that ptriv(18) = ptriv(19) = ptriv(20) = 1/2, in violation of and-drop. the same model can work as a counterexample for or-drop. let us emphasize the intuitive reason why and-drop fails. we are using a notion of probability, ptriv, which is defined relative to a domain of worlds. for some sentences a and b, it can be that the domain over which a∧b is defined is larger than the domains over which a and b are defined. in that case, it might be that the probability of a conjunction is greater than the probability of a conjunct. let us end with a note on future directions of research. it is crucial for our thesis that subjects’ likelihood judgments are boolean when conditionals are not involved. experiment 2 was designed to probe this. while it provides some evidence in favor of the boolean hypothesis, the overall results are mixed. in ongoing work, we are developing new versions of our experiments that aim at establishing the point more clearly. references bradley, richard. 2000. a preservation condition for conditionals. analysis 60(3). 219– 222. bradley, richard. 2002. indicative conditionals. erkenntnis 56(3). 345–378. 10.1023/a:1016331903927. bradley, richard. 2007. a defence of the ramsey test. mind 116(461). 1–21. cantwell, john. 2006. the laws of non-bivalent probability. logic and logical philosophy 15(2). 163–171. proceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 262 https://doi.org/10.3765/elm https://www.elm-conference.net/ charlow, nate. 2016. triviality for restrictor conditionals. noûs 50(3). 533–564. égré, paul, lorenzo rossi & jan sprenger. 2021. de finettian logics of indicative conditionals part i: trivalent semantics and validity. journal of philosophical logic 50(2). 187–213. 10.1007/s10992-020-09549-6. de finetti, bruno. 1936. la logique de la probabilité. actes du congrès international de philosophie scientifique 4. 1–9. gillies, anthony s. 2004. epistemic conditionals and conditional epistemics. noûs 38(4). 585–616. gillies, anthony s. 2009. on truth-conditions for if (but not quite only if ). philosophical review 118(3). 325–349. hájek, alan & n. hall. 1994. the hypothesis of the conditional construal of conditional probability. in ellery eells, brian skyrms & ernest w. adams (eds.), probability and conditionals: belief revision and rational decision, 75. cambridge university press. hájek, alan. 2019. interpretations of probability. in edward n. zalta (ed.), the stanford encyclopedia of philosophy, metaphysics research lab, stanford university fall 2019 edn. kaufmann, stefan. 2009. conditionals right and left: probabilities for the whole family. journal of philosophical logic 38(1). 1–53. khoo, justin. 2022. the meaning of if. new york, usa: oxford university press. khoo, justin & paolo santorio. 2018. lecture notes: probability of conditionals in modal semantics. lecture notes for a course at nasslli 2018, available at http://paolosantorio.net/ks-nasslli2018.pdf. kratzer, angelika. 1981. the notional category of modality. in h. j. eikmeyer & h. rieser (eds.), words, worlds, and contexts: new approaches to word semantics, berlin: de gruyter. kratzer, angelika. 1986. conditionals. in chicago linguistics society: papers from the parasession on pragmatics and grammatical theory, vol. 22 2, 1–15. university of chicago, chicago il: chicago linguistic society. kratzer, angelika. 2012. modals and conditionals: new and revised perspectives, vol. 36. oxford university press. lassiter, daniel. 2020. what we can learn from how trivalent conditionals avoid triviality. inquiry 63(9-10). 1087–1114. 10.1080/0020174x.2019.1698457. https://doi.org/ 10.1080/0020174x.2019.1698457. lewis, david. 1976. probabilities of conditionals and conditional probabilities. philosophical review 85(3). 297–315. rothschild, daniel. 2014. capturing the relationship between conditionals and conditional probability with a trivalent semantics. journal of applied non-classical logics 24(1-2). 144–152. 10.1080/11663081.2014.911535. russell, jeffrey sanford & john hawthorne. 2016. general dynamic triviality theorems. philosophical review 125(3). 307–339. santorio, paolo. 2022. path semantics for indicative conditionals. mind 131(521). 59–98. 10.1093/mind/fzaa101. proceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 263 https://doi.org/10.3765/elm https://www.elm-conference.net/ stalnaker, robert. 1968. a theory of conditionals. in n. recher (ed.), studies in logical theory, oxford. stalnaker, robert. 1975. indicative conditionals. philosophia 5. stalnaker, robert c. 1970. probability and conditionals. philosophy of science 37(1). 64– 80. tversky, amos & daniel kahneman. 1983. extensional versus intuitive reasoning: the conjunction fallacy in probability judgment. psychological review 90(4). 293. van fraassen, bas c. 1976. probabilities of conditionals. in foundations of probability theory, statistical inference, and statistical theories of science, 261–308. springer. proceedings of elm 2: 252-264, 2023 paolo santorio and alexis wellwood: nonboolean conditionals. 264 https://doi.org/10.3765/elm https://www.elm-conference.net/ coloring disjunction in child romanian adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bîlbîie & lyn tieu* abstract. romanian children have been shown to rarely interpret the complex disjunction sau…sau ‘either...or’ (as in sau trenul sau barca ‘either the train or the boat’) exclusively (that is, as ‘only one, not both’) in truth value judgment tasks. instead, children often favor an inclusive interpretation (‘one or both’), or a conjunctive interpretation (‘both, not just one’) (bleotu et al. 2023, 2024a). such findings contrast with those from romanian adults, who consistently interpret this disjunction exclusively. in this study, we investigate whether children interpret sau...sau more exclusively in a coloring book task (cbt), given previous evidence that children’s performance is more adult-like in tasks involving coloring rather than in truth value judgment tasks. in line with this expectation, we observed an increase in the number of exclusive-responding children compared to previous findings for romanian. however, it is important to highlight that most children still did not interpret the disjunction exclusively, indicating ongoing challenges with the interpretation of disjunction around the age of five years. keywords. first language acquisition; romanian; disjunction; exclusivity; coloring task 1. introduction. children have been argued to be more logical than adults in their interpretation of quantifiers, modals, and disjunction (see, for instance, noveck 2001, papafragou & musolino 2003). recent studies show that children’s performance on implicatures may vary with the task: while children tend to find truth value judgment tasks (tvjt) more challenging, they perform more adult-like in act-out tasks (pouscoulous et al. 2007), ternary reward tasks (katsos & bishop 2011), felicity judgment tasks (chierchia et al. 2001, foppolo et al. 2012), story-based tasks (guasti et al. 2005), and, as shown more recently, in coloring and erasing tasks (bleotu 2018, nuninga et al. 2023). in the present study, we investigate disjunction in child romanian using the coloring book task (cbt) to determine whether children’s responses align more closely with an adult-like, exclusive interpretation (that is, understanding disjunction to be true when only one of the disjuncts, and not both, is true), compared to what has been reported in tvjt studies. for instance, in response to a disjunctive utterance such as (1), adult-like children should show the * this research is supported by the project “the acquisition of disjunction in romanian” pn-iii-p1-1.1-te2021-0547 (te 140 din 30/05/2022) led by a. bleotu. a. nicolae was supported by the dfg grant ni-1850/2-1 as well as the erc synergy grant 856421 (leibnizdream). l. tieu was supported by the social sciences and humanities research council of canada and the connaught fund. a. benz’ s work was partly funded by the “linguistic meaning and bayesian modelling'” project within the leibniz collaborative excellence programme (pi anton benz, application number: k535/2023). we are grateful to the undergraduate students at the faculty of foreign languages, university of bucharest, for taking part in the experiments. we thank the children from no. 248 kindergarten in bucharest. we are also grateful to the audiences at the bucharest colloquium of language acquisition 2023 and experiments in linguistic meaning 3 (12-14 june 2024, upenn) for their useful comments and suggestions. authors: adina camelia bleotu, university of bucharest (adina.bleotu@lls.unibuc.ro; cameliableotu@gmail.com), mara panaitescu, university of bucharest (mara.panaitescu@lls.unibuc.ro), anton benz, zas berlin (benz@leibnizzas.de), andreea nicolae, zas berlin (nicolae@leibniz-zas.de), gabriela bîlbîie, university of bucharest (gabriela.bilbiie@lls.unibuc.ro), lyn tieu, university of toronto / western sydney university (marcs institute for brain, behaviour and development) / macquarie university (lyn.tieu@utoronto.ca). proceedings of elm 3: 65-74, 2025 c©2025 adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bı̂lbı̂ie, and lyn tieu published by the lsa with permission of the author(s) under a cc by license. 65 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ following behavior: if none of the objects has color, they should color just one of them, but not both; if one of these objects has color, they should do nothing; on the other hand, if both have color, they should erase the color of one of them, in line with exclusivity. (1) aş vrea să aibă culoare sau triunghiul sau cercul. would.1sg like sbjv have.sbjv.3 color or triangle.def or circle.def ‘i would like either the triangle or the circle to have color.’ 2. background. 2.1. the acquisition of disjunction. previous studies show that adults interpret disjunctive utterances containing simple disjunctions (consisting of a single disjunctive morpheme or) exclusively and inclusively, while they interpret complex disjunctions (such as either…or) predominantly exclusively in most contexts (see table 1) (chierchia et al. 2001, gualmini et al. 2001, nicolae & sauerland 2016, nicolae et al. 2024, a.o.). in contrast, children interpret both disjunctions inclusively or conjunctively (see singh et al. 2016, tieu et al. 2017), but rarely exclusively (see, nevertheless, sauerland & yatsushiro 2018, for evidence that german children can be exclusive). the hen pushed (either) the train or the boat. inclusive the hen pushed one and possibly both. exclusive the hen pushed only one, not both. conjunctive the hen pushed both, not just one. table 1: possible interpretations for the disjunctive utterance the hen pushed (either) the train or the boat children’s inclusivity is typically explained as a logical interpretation of disjunction (noveck 2001). children’s conjunctive interpretation, on the other hand, has received different explanations: (i) an implicature (singh et al. 2016, tieu et al. 2017), (ii) ambiguity between disjunction and conjunction (sauerland & yatsushiro 2018), and (iii) an experimental artifact (huang & crain 2020, skordos et al. 2020). as far as child romanian is concerned, in several tvjts, bleotu et al. (2023, 2024a) have shown romanian children to be inclusive with simplex sau ‘or’ and complex sau…sau ‘either…or’, but inclusive and conjunctive with the complex disjunction fie…fie ‘either...or’. very few children interpreted either of these disjunctions exclusively. bleotu et al. (2024b, 2024c) investigated whether children become more exclusive with a disjunctive utterance (such as the hen pushed the train or the boat) in the presence of access to the stronger conjunctive alternative (i.e., when hearing unrelated conjunctive utterances such as the deer chose a cake and a salad) or in contexts that make exclusivity relevant (such as after the non-conjunctive question did the hen push these two objects?). the findings from bleotu et al. (2024b, 2024c) suggest that neither mere access to stronger conjunctive alternatives, nor mere exposure to relevant non-conjunctive questions on their own lead to a boost in exclusive interpretations of implicatures. children are, however, more adult-like with disjunctive utterances that represent answers to conjunctive questions (such as did the hen push the train and the boat?). this suggests that both access to the conjunctive alternative and relevance of this alternative are needed to increase children’s exclusivity. these results are in contrast to those previously obtained by skordos & papafragou (2016) for quantifiers, where access to alternatives and relevance separately were found to lead to more implicatures. overall, these findings suggest that the acquisition of disjunction may be more developmentally challenging than the acquisition of quantifiers. the current study expands the proceedings of elm 3: 65-74, 2025 adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bı̂lbı̂ie, and lyn tieu: coloring disjunction in child romanian. 66 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ previous investigation and explores another potential factor that may boost exclusivity, namely the nature of the task, focusing on a special type of act-out task involving coloring items in such a way as to make the target sentences true. 2.2. act-out and coloring tasks. implicature rates in children vary significantly depending on the task. binary tvjts tend to be more challenging, while act-out tasks – involving an action such as giving rewards or moving objects to make a sentence true – yield higher implicature rates (e.g., pouscoulous et al. 2007). one argument could be that children’s varying engagement with a task can make implicatures more or less relevant in a context. another argument could be that actout tasks are preference-based tasks, and children’s preferences may be different from what they may accept when making a forced choice judgment. for instance, reward tasks have been shown to enhance implicature production. papafragou & tantalou (2004) found that children derived more implicatures when rewarding an animal based on how accurately it described its actions. similarly, katsos & bishop (2011) showed that children were sensitive to underinformativeness when giving strawberries as rewards. the authors argued that the results from binary tvjts do not reveal children’s failure with implicatures but rather their pragmatic tolerance of underinformativeness. additionally, bleotu et al. (2021a, 2021b, 2022) found that, in a reward task in which they had to reward the best descriptions with the highest reward, adults derived more implicatures from utterances containing poate ‘maybe’ than in a tvjt, and, while not fully adultlike, children also performed well in this task. pouscoulous et al. (2007) also found quite high implicature rates with french quantifiers in various age groups in an act-out task. they tested quelques ‘some’, tous/toutes ‘all’, aucun(e) ‘no’, and quelques…ne…pas ‘some…not’ utterances embedded under je voudrais ‘i would like’ in three types of scenarios: a subset scenario, where 2 of 5 boxes had tokens, an all scenario, where all 5 boxes had tokens, and a none scenario, where 0 of 5 boxes had tokens. depending on their interpretation, participants were expected to add tokens, remove tokens, or leave things as they were. for instance, in an all scenario, a puppet uttered je voudrais que quelques boîtes contiennent un jeton ‘i would like some boxes to contain a token’ in a scenario where each of five boxes already contained a token. if participants understood this utterance with a some but not all implicature, then they were expected to remove at least one token. if, on the other hand, they understood some as some and possibly all (that is, inclusively), they were expected to leave the boxes unchanged. importantly, most children here removed at least one token, thus showing evidence of having derived the implicature. given children’s relative success with quantifiers in this paradigm, we adapted it in our study to extend the investigation to disjunction. importantly though, instead of using an act-out procedure involving adding, removing, or leaving objects in place, participants were asked to color certain images, erase, or do nothing. the coloring book task (cbt) was developed by zuckerman et al. (2016) and pinto & zuckerman (2019). in this task, participants do not explicitly state their choices using words but rather indicate their preferences by coloring specific items on a coloring sheet. this methodological approach offers a notable advantage: participants do not need to verbalize their answers, but instead, their choices are reflected through their actions, which can often be more reliable than explicit verbal responses. the cbt satisfies two critical requirements for language comprehension tasks (zuckerman et al. 2016). first, it provides alternatives without explicitly presenting them. this circumvents a key methodological issue found in the tvjt, namely, the explicit presence of alternatives. according to zuckerman et al. (2016: 444), the explicit presence of alternatives forces the subject to proceedings of elm 3: 65-74, 2025 adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bı̂lbı̂ie, and lyn tieu: coloring disjunction in child romanian. 67 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ consciously consider potential interpretations that may not naturally arise, which can influence the selection process. the cbt avoids this by allowing participants to make choices through their actions rather than through explicit and metalinguistic answers. an additional advantage of the cbt, according to zuckerman et al. (2016), is its engagement appeal, particularly for children. the activity of coloring is inherently engaging, often leading to higher levels of participation. this feature makes it a useful tool for studying language comprehension in younger populations. zuckerman et al. (2016) employed the cbt to investigate (a) passives and (b) principle b of the binding theory, testing 58 dutch-speaking children aged 3;11 to 8;07, and reported more adult-like performance compared to a traditional picture selection task. similarly, gerard et al. (2017, 2018) and gerard & lidz (2018) examined children’s performance on adjunct control using both a cbt and a tvjt, and found that children provided more adult-like responses in the cbt than in the tvjt. for example, participants were presented with sentences like (2): (2) dora washed diego before pro eating the red apple. they were then asked to color the apple. depending on whether they associated the apple with dora or diego, participants’ responses reflected whether pro was co-indexed with the subject (dora) or the object (diego). figure 1: example pictures from gerard & lidz (2018) bleotu (2018) further explored the potential of the cbt by testing 18 five-year-olds on implicature derivation using quantifiers, comparing it with three other methods: the tvjt, the picture selection task, and the erasing task. the results showed that children mastered the meaning of existential quantifiers in the cbt and in the erasing task. notably, while the cbt may be argued to show that children grasped the meaning of existential quantifiers, the erasing task confirmed that they were also pragmatically sensitive to scales, deriving scalar implicatures with existential quantifiers as early as five years old. bleotu (2024) also applied the cbt to test 25 romanian-speaking five-year-olds on their understanding of modal statements, with results indicating that children understood the meaning of epistemic adverbs. nonetheless, the cbt does not always elicit adult-like responses. for instance, hall & pérez-leroux (2022) found that children’s responses were at chance when tested on comitatives (e.g., the dog with the bone is blue) but more adult-like when tested on coordinated nps (e.g., the cup and the table are green). this suggests that the cbt is able to reveal when children have a different understanding of a particular syntactic structure. thus, while the cbt generally leads to adult-like responses, particularly in the context of language comprehension, it can also highlight differences in children’s understanding, making it a valuable tool for assessing linguistic development. 3. current experiment. previous tvjt studies reveal that children are mostly inclusive with the complex disjunction sau…sau, in contrast to adults, who are typically exclusive (see bleotu et al. proceedings of elm 3: 65-74, 2025 adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bı̂lbı̂ie, and lyn tieu: coloring disjunction in child romanian. 68 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 2023, 2024a). the current study investigates whether children are (more) exclusive with the complex disjunction sau…sau in a coloring book task. 3.1. predictions. considering that children are reported to be more adult-like in act-out tasks (including coloring tasks), we expect to observe more evidence of exclusive interpretations of disjunction compared to what has previously been reported in tvjt studies (bleotu et al. 2023, 2024a). nevertheless, given previous findings that romanian children are primarily inclusive with sau…sau and often struggle with deriving exclusivity (they are more adult-like only in the presence of both explicit alternatives and relevance), children might experience challenges in deriving implicatures with disjunction more so than what has been reported for quantifiers. 3.2. participants. we tested 34 five-year-old monolingual romanian-speaking children (5;005;11, m=5;06) and 40 adult controls. 3.3. methodology. the experiment employed a coloring book task, drawing largely on the act-out task conducted by pouscoulous et al. (2007), testing children’s interpretation in multiple scenarios that required addition, removal, or no action whatsoever. participants were introduced to a puppet named bibi whose wishes they had to fulfill by coloring objects, erasing the color of objects, or taking no action. they saw displays of vehicles/fruits/shapes/vegetables in which none, some, or all of the objects were colored (figure 2a-c). figure 2: (a) 0-object scenario vs. (b) 1-object scenario vs. (c) 2-object scenario they then heard a recorded statement left by bibi on whatsapp as in (3), and they had to fulfill her wish. the materials consisted of 6 warm-up statements balanced for action (coloring/erasing/ doing nothing), 36 critical sentences, and 15 fillers (also balanced for action). the experiment tested disjunctive sentences containing the complex disjunction sau…sau, such as (3a), in three kinds of scenarios: the 0-object scenario (containing no colored objects, see figure 2a), the 1object scenario (containing one colored object, see figure 2b), and the 2-object scenario (containing two colored objects, see figure 2c). we also tested conjunctive (3b) and negative (3c) sentences as controls in these three scenarios. (3) a. bibi: aş vrea să aibă culoare sau triunghiul sau cercul. would.1sg like sbjv have.sbjv.3 color or triangle.def or circle.def ‘i would like either the triangle or the circle to have color.’ b. bibi: aş vrea să aibă culoare triunghiul și cercul. would.1sg like sbjv have.sbjv.3 color triangle.def and circle.def ‘i would like the triangle and the circle to have color.’ c. bibi: aş vrea să nu aibă culoare nici triunghiul nici would.1sg like sbjv neg have.sbjv.3 color neither triangle.def nor cercul. circle.def ‘i would like neither the triangle nor the circle to have color.’ proceedings of elm 3: 65-74, 2025 adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bı̂lbı̂ie, and lyn tieu: coloring disjunction in child romanian. 69 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3.4. predictions. for disjunctive sau…sau statements, we expected adults to color one object in the 0-object scenario, do nothing in the 1-object scenario, and erase the color of an object in the 2-object scenario, while we expected more variability in children’s answers given previously reported inclusive/conjunctive behavior (table 2). nevertheless, given the cbt’s success in eliciting adult-like performance in children, we expected some proportion of exclusive responses. scenario initial situation inclusive participants a or b, possibly (a and b) exclusive participants (a or b) but not (a and b) conjunctive participants a and b 0-object color 1 or 2 objects color 1 object color 2 objects 1-object do nothing or color 2nd object do nothing color 1 object 2-object do nothing erase 1 object do nothing table 2: predicted responses for disjunctive statements per participant type in the three scenarios 3.5. results. adults generally behaved as predicted, that is, they consistently preferred the exclusive interpretation (96.3%). turning to children, we observed generally strong performance on the conjunctive controls (89%) and the negative controls (83.3%). for the disjunctive statements, however, more non-adult-like responses were observed overall. importantly, there was variation depending on the scenario. in the 0-object scenario, 86% of children’s responses were adult-like (coloring one object); the remaining responses involved coloring two objects instead of one. in the 1-object scenario, 52.2% of responses were adult-like (doing nothing); the remaining responses involved coloring a second object. in the 2-object scenario, 44.1% of responses were adult-like (erasing the color of one object); the remaining responses involved leaving both objects colored. these results are summarized in table 3. scenario % adult-like responses 0-object 86 1-object 52.2 2-object 44.1 table 3: percentage of adult-like responses from children on target disjunctive trials an individual analysis revealed that 10/34 children were consistently exclusive, 3/34 were consistently conjunctive, and the remaining 21 children showed mixed (inclusive, and/or conjunctive, and/or exclusive) behavior. a binominal test revealed that the observed proportion of exclusive children (10/34) was significantly different from chance (p<.05). 3.6. discussion. here we discuss three points with regards to our main findings: the differences in patterns of interpretation that we observe between adults and children, as well as within children; possible explanations for the observed patterns of interpretation; and the methodological implications of our findings. first, all adults responded in a manner consistent with exclusive interpretations of the disjunctive targets. interestingly, while most children colored one object in the 0-object scenario, they varied in their behavior in the other scenarios: they would sometimes color nothing or color one more object in the 1-object scenario, and they would erase one object or simply leave the two objects colored in the 2-object scenario. thus, we observe three subgroups of children: exclusive (10 children), conjunctive (3 children), mixed (21 children). we take this behavior to suggest that some children may be at a developmental stage where they oscillate between inclusive and exclusive interpretations for the complex disjunction sau...sau, in contrast with adults, who consistently favor the exclusive interpretation. proceedings of elm 3: 65-74, 2025 adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bı̂lbı̂ie, and lyn tieu: coloring disjunction in child romanian. 70 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ based on the relatively high accuracy on the controls, we assume that children’s coloring responses essentially reflect their linguistic understanding of disjunction. that is, they deploy their semantic/pragmatic meaning for disjunction to carry out their coloring actions, to make bibi’s request true on this meaning. there are, however, two alternative non-linguistic cognitive preferences or constraints that might drive participants’ behavior in such a task, which we think can nevertheless be ruled out as explanations for the data. the first is that children might wish to color as much as possible, simply because coloring is fun (and indeed we observed that children liked coloring and even tried to use many different colors to carry out bibi’s request). the opposing strategy is to color as little as possible, in order to minimize effort. teasing apart the role of nonlinguistic preferences is difficult when they push in a similar direction as following the linguistic meaning of disjunction: in the 0-object scenario and in the 1-object scenario, having only one colored object is not only in line with inclusive/exclusive meanings of disjunction, but it is also in line with deploying minimal effort. however, in the 2-object scenario, the adult-like response (to erase the color of one object) involves both more effort and fewer colored objects than the nonadult-like response (which is to do nothing), i.e. it clashes both with deploying minimal effort and maximizing color, yet, even in this condition, a non-trivial proportion of children provided exclusive answers, i.e. they erased the color of one object. we thus argue that our results reflect children’s linguistic understanding and cannot be accounted for on non-linguistic grounds. we would like to end with the methodological implications of our study. we found that children seemed to be more adult-like with disjunction in this task compared to previous studies which used the tvjt (bleotu et al. 2023, 2024a). moreover, we found that romanian children tended to be inclusive in their interpretation of the complex disjunction sau...sau. a possible explanation for this difference could be related to the fact that the binary tvjt actually reveals that children are more pragmatically tolerant than adults (katsos & bishop 2011), rather than that they are more logical than adults. in contrast to the tvjt, the cbt is a preference-based task, which shows that at least some children in this age range prefer to interpret disjunction exclusively. while the tvjt can reveal the existence of certain interpretations, the cbt can reveal preferences for certain interpretations. the two tasks thus complement each other, shedding light on different aspects of children’s comprehension of disjunction. interestingly, there seems to be an asymmetry between preference and acceptance: children are more adult-like in their interpretive preferences than in their evaluation of the truth of sentences in context. 4. conclusion. the present findings support the use of the cbt as a method for eliciting adultlike interpretations in children. unlike the tvjt, which may simply show that children are more pragmatically tolerant than adults (katsos & bishop 2011), the cbt is a preference-based task, combining linguistic comprehension with non-linguistic production. in line with previous studies (zuckerman et al. 2016, zuckerman & pinto 2018, gerard et al. 2017, gerard et al. 2018, gerard & lidz 2018, bleotu 2018, 2019, nuninga et al. 2023), preference-based tasks like the cbt elicit more adult-like responses from children. our findings also suggest that at least some children in this age range can interpret disjunction exclusively – contra many findings from tvjt-based studies (bleotu et al. 2023, 2024a, among others), which show that romanian children are inclusive in their comprehension of the complex disjunction sau...sau. we leave for a future study a more direct comparison of children’s behavior on the cbt and a parallel tvjt, which may further shed light on how exclusivity is affected by the employment of different tasks. proceedings of elm 3: 65-74, 2025 adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bı̂lbı̂ie, and lyn tieu: coloring disjunction in child romanian. 71 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ references bleotu, adina camelia. 2018. scalar implicatures with existential quantifiers in 5-year-olds. insights from a coloring book task. presentation at the workshop scalar implicatures: formal and experimental exploration, siena, italy. bleotu, adina camelia. 2019. scalar implicatures with existential quantifiers in 5-year-olds: insights from a coloring book task and an erasing task. manuscript. bleotu, adina camelia. 2024. coloring the possible and the certain in child romanian. in exploring linguistic landscapes. a festschrift for larisa avram and andrei avram. bucureşti: bucharest university press. bleotu, adina camelia, anton benz & nicole gotzner. 2021a. where truth and optimality part. experiments on implicatures with epistemic adverbs. proceedings of experiments in linguistic meaning (elm) 1. 47–58. doi: 10.3765/elm.1.4863. bleotu, adina camelia, anton benz & nicole gotzner. 2021b. shadow-playing with romanian 5-year-olds. epistemic adverbs are a kind of magic! proceedings of experiments in linguistic meaning (elm) 1. 59–70. doi: 10.3765/elm.1.4866. bleotu, adina camelia, anton benz & nicole gotzner. 2022. romanian 5-year-olds derive global but not local implicatures with quantifiers embedded under epistemic adverbs: evidence from a shadow play paradigm. proceedings of sinn und bedeutung (sub) 26. 149–164. doi: 10.18148/sub/2022.v26i0.993. bleotu, adina camelia, rodica ivan, andreea nicolae, gabriela bîlbîie, anton benz, mara panaitescu & lyn tieu. 2023. not all complex disjunctions are alike: on inclusive and conjunctive interpretations in child romanian. proceedings of the annual conference of the cognitive science society 45. 3062–3069. bleotu, adina camelia, lyn tieu, anton benz, alexandre cremers, gabriela bîlbîie, mara panaitescu, rodica ivan, & andreea nicolae. 2024a. children interpret some disjunctions conjunctively: evidence from child romanian. psyarxiv. january 31. doi:10.31234/osf.io/bywj2. bleotu, adina camelia, gabriela bîlbîie, mara panaitescu, alexandre cremers, anton benz, andreea nicolae & lyn tieu. 2024b. does hearing and help children understand or? insights into scales and relevance from the acquisition of disjunction in child romanian. submitted. lingbuzz/008610. bleotu, adina camelia, andreea nicolae, anton benz, gabriela bîlbîie, mara panaitescu & lyn tieu. 2024c. does relevance without explicit alternatives boost exclusivity implicatures of disjunction? to appear in proceedings of the west coast conference of formal linguistics (wccfl) 42. chierchia, gennaro, stephen crain, maria teresa guasti, andrea gualmini & luisa meroni. 2001. the acquisition of disjunction: evidence for a grammatical view of scalar implicatures. in amy h.-j. do, laura dominguez & anders johansen (eds.), proceedings of the 25th boston university child language development conference (bucld), 157–168. somerville, ma: cascadilla. foppolo, francesca, maria teresa guasti & gennaro chierchia. 2012. scalar implicatures in child language: give children a chance. language learning and development 8. 365–394. gerard, juliana, jeff lidz, shalom zuckerman & manuela pinto. 2017. similarity-based interference and the acquisition of adjunct control. frontiers in psychology 8. 1822. doi: 10.3389/fpsyg.2017.01822. proceedings of elm 3: 65-74, 2025 adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bı̂lbı̂ie, and lyn tieu: coloring disjunction in child romanian. 72 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ gerard, juliana & jeff lidz. 2018. before and after the acquisition of adjunct control. in anne b. bertolini & maxwell j. kaplan (eds.), proceedings of the 42nd annual boston university conference on language development (bucld), 266–279. somerville, ma: cascadilla press. gerard, juliana, jeff lidz, shalom zuckerman & manuela pinto. 2018. the acquisition of adjunct control is colored by the task. glossa: a journal of general linguistics 3(1). 1–22. doi: 10.5334/gjgl.547. gualmini, andrea, stephen crain, luisa meroni, gennaro chierchia & maria teresa guasti. 2001. at the semantics/pragmatics interface in child language. proceedings of semantics and linguistic theory (salt) 11. 231–247. guasti, maria teresa, gennaro chierchia, stephen crain, francesca foppolo, andrea gualmini & luisa meroni. 2005. why children and adults sometimes (but not always) compute implicatures. language and cognitive processes 20(5). 667–696. doi: 10.1080/01690960444000250. hall, erin & anna pérez-leroux. 2022. children’s comprehension of np embedding. glossa: a journal of general linguistics 7(1). 1–41. doi: 10.16995/glossa.5816. huang, haiquan & stephen crain. 2020. when or is assigned a conjunctive inference in child language. language acquisition 27(1). 74–97. doi:10.1080/10489223.2019.1659273. katsos, napoleon & dorothy v.m. bishop. 2011. pragmatic tolerance: implications for the acquisition of informativeness and implicature. cognition 120(1). 67–81. doi: 10.1016/j.cognition.2011.02.015. nicolae, andreea & uli sauerland. 2016. a contest of strength: or versus either–or. in polina berezovskaya nadine bade & anthea schöller (eds.), proceedings of sinn und bedeutung (sub), vol. 20, 551–568. open journal systems. nicolae, andreea, aliona petrenco, anastasia tsilia & paul marty. 2024. exclusivity and exhaustivity of disjunction(s): a cross-linguistic study. to appear in proceedings of sinn und bedeutung (sub) 28. noveck, ira. 2001. when children are more logical than adults. cognition 78(2). 165–188. doi: 10.1016/s0010-0277(00)00114-1. nuninga, rosa, ileana grama, jeannette schaeffer, charlotte jurrien, manuela pinto & shalom zuckerman. 2023. deriving scalar and ad-hoc implicatures in an ecologically valid task. presentation at the conference bucharest colloquium of language acquisition 8, workshop logical operators: theory and acquisition. papafragou, anna & julien musolino. 2003. scalar implicatures: experiments at the semantics pragmatics interface. cognition 86(3). 253–282. doi: 10.1016/s0010-0277(02)00179-8. papafragou, anna & niki tantalou. 2004. children’s computation of implicatures. language acquisition 12(1). 71–82. doi: 10.1207/s15327817la1201_3. pinto, manuela & shalom zuckerman. 2019. coloring book: a new method for testing language comprehension. behavior research methods 51(6). 2609-2628. pouscoulous, nausicaa, ira noveck, guy politzer & anne bastide. 2007. a developmental investigation of processing costs in implicature production. language acquisition 14. 347– 375. doi: 10.1080/10489220701600457. r core team. 2021. r: a language and environment for statistical computing. r foundation for statistical computing. vienna, austria. available online at https://www.r-project.org/. sauerland, uli & kazuko yatsushiro. 2018. the acquisition of disjunctions: evidence from german children. proceedings of sinn und bedeutung (sub) 21(2). 1065–1072. proceedings of elm 3: 65-74, 2025 adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bı̂lbı̂ie, and lyn tieu: coloring disjunction in child romanian. 73 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ singh, raj, ken wexler, andrea astle-rahim, deepthi kamawar & danny fox. 2016. children interpret disjunction as conjunction: consequences for theories of implicature and child development. natural language semantics 24(4). 305–352. doi: 10.1007/s11050-016-91263. skordos, dimitrios & anna papafragou. 2016. children’s derivation of scalar implicatures: alternatives and relevance. cognition 153. 6–18. skordos, dimitrios, roman feiman, alan c. bale & david barner. 2020. do children interpret “or” conjunctively? journal of semantics 37(2). 247–267. tieu, lyn, kazuko yatsushiro, alexandre cremers, jacopo romoli, uli sauerland & emmanuel chemla. 2017. on the role of alternatives in the acquisition of simple and complex disjunctions in french and japanese. journal of semantics 34(1). 127–152. doi: 10.1093/jos/ffw010. zuckerman, shalom, manuela pinto, elly koutamanis & yoïn van spijk. 2016. a new method for language comprehension reveals better performance on passive and principle b constructions. in jennifer scott & deb waughtal (eds.), proceedings of the 40th annual boston university conference on language development (bucld), 443–456. somerville, ma: cascadilla press. zuckerman, shalom & manuela pinto. 2018. age of acquisition ratings validated by actual vocabulary scores. poster presented at architectures and mechanisms for language processing (amlap), berlin, germany. zuckerman, shalom & manuela pinto. 2020. the acquisition of ‘bridging’ tested with the coloring book method. in pedro guijarro-fuentes & cristina suárez-gómez (eds.), new trends in language acquisition within the generative perspective, 289–311. dordrecht: springer. proceedings of elm 3: 65-74, 2025 adina camelia bleotu, mara panaitescu, anton benz, andreea nicolae, gabriela bı̂lbı̂ie, and lyn tieu: coloring disjunction in child romanian. 74 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ using the lower bound set by the universal modal to investigate the status of partial objects and count nouns premvanti patel, kristen syrett & athulya aravind* abstract. prior research has demonstrated that when given objects (e.g., forks) broken into pieces, children deviate from adults by counting each discrete object-piece as on par with a whole. a recent proposal ties this behavior to the vagueness and contextsensitivity inherent to count noun semantics. the present study leverages the universal modal have to in order to investigate how a linguistic context, one which sets lower bounds on numerals in its scope, regulates nominal application. our results show that for children, who prefer the ‘exact’ reading of numerals, the partial object not only serves to meet the lower bound, but also exceeds a numerical upper bound. adults, on the other hand, do not consider the partial object as meeting the lower bound induced by the modal. because we cannot determine the explanation for this finding with our current design, we plan to adapt it to use the existential modal allowed to. keywords. partial objects; modals; numerals; gradability; context-sensitivity; count noun semantics 1. introduction. count nouns such as ball and fork are among the first words children produce. yet, children show a surprising, non-adult-like willingness to apply such words not just to whole balls and forks, but also to their discrete parts. in a seminal study, shipley & shepperson (1990) gave children sets of whole objects and object parts (“partial objects”) with specific instructions of what to count. when shown a set as in figure 1, with four whole forks and two broken forkpieces, and asked to “count the forks”, children tended to count 6, as if treating the partial objects on par with the wholes. figure 1: example counting prompt from shipley & shepperson (1990) * we are grateful for all participating families and individuals, and to maya brisman, destiny eversole, denley kofoed, margaret morgan, divya natarajan, jamie oliver, david phan, shuyan wang, and eunice zhang for their help with data collection. we would also like to thank kate kinnaird and indira das for their support in coordinating this collaboration. authors: premvanti patel (pspatel@mit.edu) and athulya aravind (aaravind@mit.edu), massachusetts institute of technology & kristen syrett (kristen.syrett@rutgers.edu), rutgers university. proceedings of elm 3: 289-298, 2025 c©2025 premvanti patel, kristen syrett, and athulya aravind published by the lsa with permission of the author(s) under a cc by license. 289 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ adults either counted 4, ignoring the partial objects entirely, or 5, combining the two forkparts into an imagined whole and counting it as 1. the authors took these findings to indicate that children have a conceptual bias to treat discrete physical objects as “default units” for counting. since then, many have replicated and extended shipley and shepperson’s original findings, and proposed different accounts of children’s part-counting behavior. the details of these theories vary substantially: some argue that the problem lies in children's still-developing conceptual knowledge (wagner & carey 2003), others suggest it stems from the interaction between cognitive defaults and developing noun semantics (brooks et al. 2011), and some point to children’s more limited access to lexical alternatives to refer to parts of objects or degraded objects (srinivasan et al. 2013). despite these differences, most of these theories converge on the idea that children's behavior fundamentally differs from how adults handle partial objects. in contrast, a more recent account by syrett & aravind (2022) proposed that children’s performance may be consistent with an adult noun semantics at its core. count nouns, for both adults and children, have meanings that depend on context to determine what counts as a unit (krifka 1989, rothstein 2010), and in certain contexts, either a whole or a partial object could fall under the extension of the count noun. this idea gains support from anecdotal examples of adults using count nouns to label partial objects similarly as they would whole objects—for example, when presenting findings from an archeological dig, a shard of pottery could be described as a ‘plate’ tout court. what differentiates adults from children, they argue, is adults’ more sophisticated ability to integrate context-specific information to restrict the noun’s application on a case-by-case basis. to test their context-sensitivity hypothesis, syrett & aravind presented the participants with a task in which they had to determine whether a partial object counted as an instance of a count noun like ball, given the presence or absence of a speaker goal related to the object. children, but not adults, were influenced by the degree of contextual support in deciding whether a partial object was a suitable referent for a count noun. when children were told, for example, that someone intended to play tennis with a ball, they were less likely to accept a partial object as a referent for “a ball.” these results demonstrate that there are limits to children’s application of count nouns to partial objects, modulated by contextual information. but whereas syrett & aravind’s task probed object reference (and therefore, category membership as indicated by the count noun and what falls under its extension), much of the prior work was focused on counting and quantification of sets of objects. our goal in this paper is to test the context-sensitivity account in a task that calls upon participants to quantify objects, without overtly counting them. to achieve this, we leverage the bounding conditions induced by modals and their effect on the interpretation of numerical expressions to determine the status of partial objects relative to wholes. if children malleably treat partial objects as either parts of wholes or wholes depending on the context, then in a quantification task where the goal is to meet and exceed a lower bound, children should allow partial objects to serve this purpose. adults, however, should continue to recognize partial objects as such, and not allow them to meet the lower bound. previewing our results, we find that the adult-child difference in how partial objects are treated re-emerges in this task. 2. bounding conditions on numerals in the scope of modals. depending on the environment in which they appear, numerals can receive different interpretations: an ‘exact’, ‘at least’ or ‘at most’ reading. in a discourse context such as (1), the numeral three is naturally understood to mean ‘exactly three’—as in, providing both upper and lower limits on the number of mistakes made. if proceedings of elm 3: 289-298, 2025 premvanti patel, kristen syrett, and athulya aravind: using the lower bound set by the universal modal to investigate the status of partial objects and count nouns. 290 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ b made exactly three mistakes, answering a’s question with any other numeral would be misleading, or even untruthful. (1) a: how many mistakes did you make? b: i made three mistakes. in a somewhat different context, such as (2), the numeral is understood as providing a lower bound meaning ‘at least three’. b’s response can be understood as felicitous and truthful even if they, in fact, have four children. (2) a: people with three children get a tax break. do you have three children? b: yes, i have three children. the surrounding linguistic context can highlight different readings of numerals. in the scope of a universal modal such as have to, numerals are most naturally understood as having a lowerbounded, ‘at least’, reading. thus, (3) is understood as implying that anyone with three or more children will receive the tax break. numerals in the scope of an existential modal like allowed to, in contrast, typically receive an upper-bounded reading. thus, (4) implies that anyone who makes three or fewer mistakes will pass the test. (3) you have to have three children to receive the tax break. (4) you are allowed to make three mistakes and still pass the test. prior developmental work suggests that children can access these different interpretations of numerals, despite sometimes showing less flexibility than adults in their interpretations. in unembedded contexts, children have been shown to prefer an ‘exact’ interpretation of numerals (e.g., huang & snedeker 2009, huang et al. 2013, papafragou & musolino 2003). when the numeral is in the scope of modals, however, children more readily access the upperand lowerbounded readings. musolino (2004) used a truth-value judgment task to test fourand five-year-olds’ understanding of the interaction between numerals and modals by pairing stories with various modal statements that induced either a lower bound in the ‘at least’ condition, as in (5), or an upper bound in the ‘at most’ condition, as in (6). in both cases, children were asked, if the troll won the coin. (5) goofy said that the troll had to put two hoops on the pole in order to win the coin. (6) goofy said the troll could miss two hoops and still win the coin. initially, given the contexts above, while children were able to access the upper-bounded readings of numerals in the ‘at most’ condition, they struggled with lower-bounded readings in the ‘at least’ condition and favored ‘exact’ readings. however, in a second experiment, which featured simplified scenarios and modal statements such as (7), children much more readily accessed lowerbounded readings of numerals. (7) let’s see if goofy can help the troll. the troll needs two cookies. does goofy have two cookies? kennedy & syrett (2022) expanded on musolino’s experiments by adding a condition missing from musolino’s design: one in which the upper bound, induced by the existential modal allowed to, was exceeded. across trials, and within participants, they manipulated the quantity of objects so that a character took less than 2, exactly 2, or greater than 2 objects or measurements of substances. their scenarios were paired with modal statements, such as (8). proceedings of elm 3: 289-298, 2025 premvanti patel, kristen syrett, and athulya aravind: using the lower bound set by the universal modal to investigate the status of partial objects and count nouns. 291 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (8) you are allowed to use two lemons. children successfully rejected the action when the upper bound was exceeded, and were also significantly less likely to accept actions that did not meet the upper limit. these studies provide a good starting point for our experiment, which assesses children’s and adults’ treatment of partial objects and the role of language in mediating how such objects are quantified. here, we focus on the universal modal have to, which induces a lower bound on numerals in its scope (as in (3)), to investigate how a linguistic context affects children’s or adults’ numerical judgments. we ask whether partial objects serve to satisfy the numerical lower bound in a modal statement such as (3). if they can be both flexibly treated as category members denoted by the count noun in the right contexts and also counted as one unit, partial objects should count towards meeting the limit. our question is whether this is the case for both children and adults, in a task focused on quantification. 3. experiment. sample size, procedures, and analyses for this experiment were pre-registered at https://osf.io/phyds. 3.1. participants. 40 english-acquiring children (4;6-5;6, m=4;11) and 21 english-speaking adults (n=40 preregistered, in prog.) participated in the experiment. an additional 10 children were tested but excluded from the full sample for failing comprehension checks (4), not completing the experiment (3), inattentiveness (1), less than 50% home exposure to english (1), or experimenter error (1). all children were recruited from a database of families interested in participating in research with the mit language acquisition lab, and zoom-tested by a live experimenter. four additional adults were tested as well, but excluded from the full sample for failing comprehension checks (2) or not completing the experiment (2). all adults were undergraduate students from rutgers university who received extra credit in a linguistics or cognitive science course for their participation, and took an asynchronous variant of the child experiment, in which video clips of each trial were inserted into a self-paced qualtrics survey. 3.2. materials. participants were introduced to a game in which characters had to satisfy a rule to receive a reward. across trials, object kind and the distribution of whole and partial objects in the sets shown to participants varied, while the numeral in the instruction sentences was always ‘three’. participants gave a star when they found the numerical conditions to be met, and a calculator (or ‘counting machine’) otherwise. see figure 2. proceedings of elm 3: 289-298, 2025 premvanti patel, kristen syrett, and athulya aravind: using the lower bound set by the universal modal to investigate the status of partial objects and count nouns. 292 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 2: (a) during training, participants were shown a trial in which the pictured object did not match the noun in the modal statement, and one in which it did. catch items were designed similarly. (b) test trials split by set type (whole-only and whole+partial), all for which participants heard the same pre-recorded modal statement. all participants saw two types of trials based on the nature of the object sets on display. the ‘whole-only’ set type included trials with 2, 3, or 4 identical whole objects. the ‘whole+partial’ set type consisted of object sets with either 2 or 3 identical whole objects, plus a single partial object. the partial object stimuli were created by removing portions from the images of the whole objects, but were identical in all other respects. altogether, there were object sets with the following five cardinalities: 2, 2.5, 3, 3.5, 4. each object set was paired with a modal statement such as (9), in which we used count nouns denoting everyday objects (e.g., balls, cups, forks). (9) to get a star, you have to have three forks. 3.3. procedures. all participants saw trials in both the whole-only and whole+partial set types in a pseudorandomized order. the experimental session began with an introduction to zoryn, the protagonist, and her friends, a group of aliens visiting earth. participants were invited to play a counting game with them, in which they had to listen to a rule provided by zoryn and then determine whether the set of objects in a friend’s possession was rule-compliant. if they followed the rule, they should be rewarded with a star; if not, they should receive a calculator to count better the next time. to ensure that participants understood the task, they saw two training trials prior to the experimental phase. the instruction sentences for these trials involved the universal modal, but without a numeral (e.g., “to get a star, you have to have a banana”). in one trial, the friend had an object that matched the noun in the rule, and in the other, they had one that did not. the test phase consisted of 18 total trials: 2 per cardinality for the whole-only trials (6 total), 4 per cardinality for the whole+partial trials (8 total), and 4 catch items. these catch items followed a similar structure as the training trials, and served as both task-comprehension and attention checks. for each trial, we coded whether participants judged a set as compliant with the proceedings of elm 3: 289-298, 2025 premvanti patel, kristen syrett, and athulya aravind: using the lower bound set by the universal modal to investigate the status of partial objects and count nouns. 293 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ modal statement, with a choice of star indicating ‘yes’ (coded as 1) and a choice of calculator indicating ‘no’ (coded as 0). 3.4. predictions. for the whole-only trial types (2-object, 3-object, and 4-object), we expect that all participants, regardless of their interpretation of the numeral, will agree on 2-object and 3object trials, but vary on 4-object trials based on how they interpret the phrase have to have three n. in 2-object trials, participants should respond ‘no’, as the lower bound is not met. on the other hand, in 3-object trials, the lower bound is met, so participants are expected to respond ‘yes’. in 4-object trials, participants who prefer an ‘at least’ reading of the numeral should respond ‘yes’ (since the lower bound is met and exceeded). thus, for the same reason, those who prefer an ‘exact’ reading should say ‘no’, as not only is the lower bound met, but the upper bound is exceeded. this pattern might be observed with children, who have previously been shown to most readily access ‘exact’ readings of numerals (huang, spelke, & snedeker, 2013; papafragou & musolino, 2003). for the whole+partial trials, expectations vary based on two factors: how participants quantify partial objects based on context, and their preferred reading of the numeral. if partial objects can count as 1 unit as the context demands, then sets with 2 whole objects and 1 partial object could be treated as on par with a set containing 3 wholes, leading to a ‘yes’ response in the 2.5-trials. on the other hand, if the partial objects do not count as having a cardinality of 1, 2.5trials should yield ‘no’ responses, in contrast to 3-trials. as for the numeral, if participants access a lower bounded reading, 3.5-trials should be accepted irrespective of their treatment of the partial object. if they access only an ‘exact’ reading, 3.5-trials may yield ‘no’ responses if the partial object counts as an instance of the noun, pushing the set beyond the limit. see figure 3. figure 3: predictions for behavioral responses from participants in response to the alien’s selection of object quantities, given the rule you have to have three forks, where a star corresponds to a ‘yes’-response (coded as ‘1’), and a calculator corresponds to a ‘no’-response (coded as ‘0’). 3.5. analyses. all analyses were performed using r statistical software (v4.3.1; r core team 2023). our primary question was how the rates of accepting a set as rule-compliant vary based on set composition, and whether this differs across adult and child populations. to test these questions, we fit separate logistic mixed effects models for the whole-only and whole+partial set types. for each, we predicted the probability of responding ‘yes’ as a function of set type and age proceedings of elm 3: 289-298, 2025 premvanti patel, kristen syrett, and athulya aravind: using the lower bound set by the universal modal to investigate the status of partial objects and count nouns. 294 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ group, with random intercepts for participants1. the model syntax was as follows (where ‘settype’ refers to the number of objects in the set): yesresponse ~ settype * agegroup + (1|participant). for the whole-only model, set type had three levels (2, 3, 4). for the whole+partials model, set type had two levels (2.5, 3.5). age group had two levels (adults, children) in both models. all factors were treatment-coded. to test for main effects and interactions, we used log-likelihood chisquared tests to compare models with and without the relevant effect. figure 4: mean ‘yes’-responses (+/1 sem) split by set type and age group for the whole-only trial types (fig.4, left panel), model comparisons revealed that age group (χ2(1)=4.1, p = 0.04), set type (χ2(2) = 48.2, p < .001), and their interaction (χ2(2) = 9.3, p = 0.001) significantly improved model fit. the age group effect is driven by adults being more likely to respond ‘yes’ than children. participants in both groups were also more likely to respond ‘yes’ to 3-object trials—corresponding to an ‘exact’ reading of the numeral in the rule. post-hoc tests exploring the interaction reveal that while both populations were alike in their judgments of 2-object and 3-object trials, they diverge in their responses to 4-object trials: adults were significantly more likely than children to respond ‘yes’ to 4-object trials (β=-2.5, se= 0.7, z=3.5 p <.001), suggesting that they were able to access the ‘at least’ reading, which children resisted. in the whole+partial trial types (fig. 4, right panel), children and adults also differed in their judgments. model comparisons revealed a significant interaction of age group and set type (χ2(1)= 73.7, p<0.001). this interaction was driven by the fact that children’s and adults’ ‘yes’-responses patterned in opposite directions for 2.5-object and 3.5-object trials: children were significantly more likely than adults to accept sets of 2.5 objects as rule-compliant and meeting the lower bound of 3 (β=3.9, se=0.6, z=6.3, p<0.001). at the same time, they were significantly less likely than adults to do so for sets of 3.5 objects (β=-1.8, se=0.6, z=-3.4, p=0.001). in other words, children saw a set containing 2 wholes and 1 partial object as satisfying the bounds 1 using more complex random effect structures led to problems with model convergence. proceedings of elm 3: 289-298, 2025 premvanti patel, kristen syrett, and athulya aravind: using the lower bound set by the universal modal to investigate the status of partial objects and count nouns. 295 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ conditions set by “have to have three n”, whereas a set containing 3 wholes and 1 partial object exceeded it. 4. discussion. in our study, we used a quantification task probing numerical judgments to better understand children’s and adults’ treatment of partial objects. we asked whether participants’ willingness to treat a partial fork as falling under the extension of a count noun like fork, and counting it as one fork-unit, depended on contextual factors as previously proposed (e.g., syrett & aravind 2022). if so, can information contained in the sentence itself modulate whether a partial object counts as 1? to address these questions, we capitalized on the interaction of numerals and the universal modal have to, which favors a lower-bounded, or ‘at least’, reading of a numeral in its scope. specifically, we asked, when given a rule such as “you have to have three forks” paired with a set of 2 whole and 1 partial forks, are participants willing to count the partial fork as 1 if doing so serves to meet the lower bound conditions set by the modal? our results answer this question in interestingly different ways for our two populations. adults were able to access the lower-bounded, ‘at least’ interpretation of the numeral, as evidenced by their acceptance of sets containing 4 whole objects. however, they rarely accepted sets containing two wholes and one partial object. this result indicates that for adults, the partial object was treated on par with a whole, so as to satisfy the lower bound. in order to gain further insight into how adults (and children) quantify partial objects in light of an interaction between a modal and a numeral, we are currently conducting a complementary experiment that features the existential modal allowed to, which induces an upper bound on a numeral in its scope, as in (10) below and (4) above. (10) employees are allowed to take three snacks from the break room. in these examples, the numeral three receives an upper-bounded, ‘at most’, interpretation. while taking up to three snacks (or making up to three mistakes) would be acceptable, more than three would exceed the upper bound and be unacceptable relative to the conditions imposed by the modal. this experiment will allow us to determine if a partial object can serve to exceed the upper bound and incur a penalty, even while for adults, it does not serve to meet a lower bound. children diverged from adults in having a strong preference for the ‘exact’ reading of the numeral, demonstrated by their low acceptance of sets of 4. likewise, with sets containing three whole objects and one partial object, children responded ‘no’. thus, for them, the partial object not only served to satisfy the numerical lower bound, but also served to push the cardinality of a set beyond the upper bound. they also accepted sets with two whole objects and one partial object as meeting the lower bound almost as often as they did sets of three whole objects. this pattern highlights a second, critical divergence from adults: for children, a partial object does count as one instance of the relevant count noun, thereby meeting the lower bound of three. this finding is in line with earlier work showing that in tasks involving counting and quantification, children treat partial objects on par with wholes. our results raise key questions about the source of the current results, especially against the backdrop of previous tasks. in many quantification-focused tasks, children have consistently treated partial objects as wholes. and yet, in syrett & aravind (2022), they resisted this treatment when given contextual factors (there, a speaker-articulated goal depending on the object performing a function dependent on its shape). it may be that when the focus is on quantification of objects, children treat partial objects as wholes, in the absence of a more fine-grained scalar system allowing for partial objects to be measured as fractional portions (which adults have). as a result, children gravitate to quantities of 0 and 1, without a middle ground—even as their proceedings of elm 3: 289-298, 2025 premvanti patel, kristen syrett, and athulya aravind: using the lower bound set by the universal modal to investigate the status of partial objects and count nouns. 296 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ linguistic production highlights halves, pieces, and broken parts (as they did in their justifications for us here). it may instead be that children were sensitive to the context of this experiment, which privileged a requirement to meet lower bounds, and their sensitivity to this contextual goal influenced their treatment of partial objects as wholes satisfying this lower bound. both explanations allow for the possibility that for both children and adults, noun application and object individuation may diverge; that is, the question of whether an object can serve as a referent for a noun and how to count or measure that entity may not always align. we find the comparison between count noun application and nominal reference on the one hand, and object individuation and quantification on the other, to be an exciting avenue to pursue. although adults did not treat the partial object as meeting the lower bound, we cannot say that they excluded it from their counts entirely. there is, in fact, suggestive evidence that this is not the case. had adults ignored the partial object altogether, we would expect their treatment of the 3.5 set to be comparable to their treatment of the 3 set: the 3.5 set contains 3 whole objects and an irrelevant object that doesn’t count. but numerically, adults’ yes-responses to the 3.5 set were comparable to their treatment of the 4 set, and lower than their yes-responses to the 3 set. in other words, for adults, the 3.5 set does not seem to satisfy, the ‘exact’ reading of the numeral. references brooks, neon, amanda pogue & david barner. 2011. piecing together numerical language: children’s use of default units in early counting and quantification. developmental science 14(1). 44–57. https://doi.org/10.1111/j.1467-7687.2010.00954.x. huang, yi ting & jesse snedeker. 2009. semantic meaning and pragmatic interpretation in 5yearolds: evidence from real-time spoken language comprehension. developmental psychology 45(6). 1723–1739. https://doi.org/10.1037/a0016704. huang, yi ting, elizabeth spelke & jesse snedeker. 2013. what exactly do numbers mean? language learning and development 9(2). 105–129. https://doi.org/10.1080/15475441.2012.658731. kennedy, christopher & kristen syrett. 2022. numerals denote degree quantifiers: evidence from child language. in nicole gotzner & eli sauerland (eds.), measurements, numerals and scales, 135–162. cham: palgrave macmillan. krifka, manfred. 1989. nominal reference, temporal constitution and quantification in event semantics. in renate bartsch, johan van benthem & peter van emde boas (eds.), semantics and contextual expression, 75–116. mouton: de gruyter. musolino, julien. 2004. the semantics and acquisition of number words: integrating linguistic and developmental perspectives. cognition 93(1). 1–41. https://doi.org/10.1016/j.cognition.2003.10.002. papafragou, anna & julien musolino. 2003. scalar implicatures: experiments at the semantics– pragmatics interface. cognition 86(3). 253–282. https://doi.org/10.1016/s0010 0277(02)00179-8. r core team. 2021. r: a language and environment for statistical computing. r foundation for statistical computing, https://www.r-project.org/. rothstein, susan. 2010. counting and the mass/count distinction. journal of semantics 27(3). 343–397. https://doi.org/10.1093/jos/ffq007. shipley, elizabeth f. & barbara shepperson. 1990. countable entities: developmental changes. cognition 34(2). 109–136. https://doi.org/10.1016/0010-0277(90)90041-h. proceedings of elm 3: 289-298, 2025 premvanti patel, kristen syrett, and athulya aravind: using the lower bound set by the universal modal to investigate the status of partial objects and count nouns. 297 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ srinivasan, mahesh, eleanor chestnut, peggy li & david barner. 2013. sortal concepts and pragmatic inference in children’s early quantification of objects. cognitive psychology 66(3). 302–326. https://doi.org/10.1016/j.cogpsych.2013.01.003. syrett, kristen & athulya aravind. 2022. context sensitivity and the semantics of count nouns in the evaluation of partial objects by children and adults. journal of child language 49(2). 239–265. https://doi.org/10.1017/s0305000921000027. wagner, laura & carey, susan. 2003. individuation of objects and events: a developmental study. cognition 90(2), 163–191. https://doi.org/10.1016/s0010-0277(03)00143-4 proceedings of elm 3: 289-298, 2025 premvanti patel, kristen syrett, and athulya aravind: using the lower bound set by the universal modal to investigate the status of partial objects and count nouns. 298 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ social meaning and pragmatic reasoning: the case of (im)precision stephanie solt, roland mühlenbernd, & mariya burbelko* abstract. on the basis of a speaker’s choice between linguistic alternatives, a hearer can draw inferences not only about facts of the world, but also about the social properties of the speaker. the goal of this work is to investigate how such social meanings arise, particularly in the case where the alternatives in question differ in their core logical or semantic meaning. taking variation in numerical precision level as a case study, we seek to test the broad general hypothesis that social inferences may be derived via pragmatic reasoning about the needs of the situation, the epistemic state of the speaker, and the reasons for their choice of form. we report on two matched guise studies which demonstrate that the social meaning of (im)precision is sensitive to context and (to some extent) speaker knowledge state, and are correlated with inferred reasons for expression choice, findings which support the predictions of the pragmatic view. keywords. social meaning; alternatives; gricean reasoning; numerical expressions; approximation; matched guise methodology 1. introduction. a speaker’s choice between linguistic alternatives can prompt their hearer to draw pragmatic inferences about facts of the world (e.g. the inference from the utterance of some that all does not obtain). but such choices can also invite inferences about the properties, ideologies and/or stances of the speaker herself; that is, they can convey social meaning. recently, there has been growing interest in exploring the connections between social meaning (traditionally studied within sociolinguistics; e.g. eckert 2012) and pragmatic reasoning and processes (burnett 2019 on sociophonetic variation, acton 2019 on the definite article, beltrama & papafragou 2023 on relevance and informativity). in the present work, we investigate this topic from the perspective of the phenomenon of numerical (im)precision, that is, the choice of the level of granularity at which numerical information is reported. for example: (1) a. the train departs at 8:03. precise b. the train departs at around 8 o’clock. approximate previous work has shown that the choice of precision level can convey social meaning (beltrama 2018, beltrama et al. 2022; the latter henceforth bsb). speakers who use precise forms are perceived as more intelligent/articulate/confident (statusor competence-related) than those who use approximate forms, but also more pedantic/uptight, while those who use approximate forms are seen as more likeable/friendly/laidback (solidarityor likeability-related). the goal of the present research is to shed light on how these inferences arise. in the case of sociophonetic variation such *funded by the deutsche forschungsgemeinschaft (dfg, german research foundation) – sfb 1412, 416591334. for helpful comments, we would like to thank andrea beltrama, heather burnett, uli sauerland, and the audiences at the zas, the llc at university of paris, and elm3. authors: stephanie solt, leibniz-zentrum allgemeine sprachwissenschaft (solt@leibniz-zas.de), roland mühlenbernd, leibniz-zentrum allgemeine sprachwissenschaft (muehlenbernd@leibniz-zas.de) & mariya burbelko, leibniz-zentrum allgemeine sprachwissenschaft and hu berlin (burbelko@leibniz-zas.de). proceedings of elm 3: 371-382, 2025 c©2025 stephanie solt, roland mühlenbernd, and mariya burbelko published by the lsa with permission of the author(s) under a cc by license. 371 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ as that involving the velar and alveolar realizations of the variable ing, as in working vs. workin’ (campbell-kibler 2007, 2011), the usual view is that social meanings are indexically associated with the linguistic forms themselves. but this is less plausible when the alternatives in question differ in their semantic content, as is the case for alternations such as (1). focusing on this case study, we seek to test the following broad hypothesis: the social meaning of (im)precision is derived via pragmatic reasoning about the needs of the situation, the epistemic state of the speaker, and the reasons for their choice of form. specifically, we hypothesize that the competence-related associations of precise forms derive from the inference that the speaker knows the exact value (i.e. has a high knowledge level), something that is not necessarily the case for a speaker who uses an approximation. conversely, the likeability-related associations of approximate forms are hypothesized to derive from the inference that the speaker, in a situation where high precision is not required, has chosen to ‘round off’ the reported value to make the information easier for their hearer to understand (van der henst et al. 2002, solt et al. 2017). finally, the association of precise forms with pedantry derives from the inference that the speaker is being more precise than required in the utterance situation, highlighting their knowledge and not engaging in hearer-oriented simplification. this pragmatic view of social meaning leads to the following three predictions: prediction 1: context dependence the observed social meaning of (im)precision will be modulated by the utterance context, in particular the degree of precision required: the competencerelated associations of precise forms will be most pronounced in a situation where high precision is required (e.g. making a police report), whereas the likeability-related associations of approximate forms and pedantry-related disadvantages of precise forms will be most pronounced in contexts where high precision is not required (e.g. a casual chat with friends). bsb found certain contextual effects of this nature, but these were not entirely robust; this may relate to the complexity of the study design (12 conditions), but also to the fact that the tested scenarios could not be directly linked to contextual precision needs. prediction 2: sensitivity to available information if it is established in the context that the speaker has the precise information available to consult (for example, in the form of a train schedule in (1)), then the competence-related associations of high precision will be diminished (since simply reading out available information does not demonstrate a high level of knowledge), whereas the likeability-related associations of approximation in low-precision contexts will be strengthened (since it can be more reliably inferred that the speaker is choosing to speak approximately for heareror situation-based reasons). prediction 3: correlation with motivations the social meaning of alternative numerical forms will be correlated with the motivation attributed to the speaker for their choice of form. in particular, if the perceived motivation for the use of an approximate form is the lack of precise knowledge, this is expected to correlate with lower competence ratings, whereas if it is desire to make the information easier to understand, this is expected to correlate with higher likeability ratings. we test these predictions in two pre-registered experiments.1 1the full preregistration can be found at https://osf.io/r34vt. proceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 372 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 2. pretest. as a first step, 16 scenarios were created in which a speaker asks a question requiring a numerical answer; each had 2 versions, one expected to prefer a precise answer (highpr), the second expected to prefer an approximate answer (lowpr). these were tested in an online experiment in which self-reported native speakers of english with u.s. and u.k. ip addresses (n=80) were recruited via prolific and paid £1.05 for participation. each participant saw 8 of the 16 scenarios (randomly assigned), each in one of its two versions (also randomly assigned), as well as two possible answers (precise, approximate), and were asked to indicate which of the two answers was more appropriate, or if both were equally appropriate. an additional 8 filler items were included. based on responses to selected filler items serving as attention checks, 6 participants were excluded, leaving 74 participants for the analysis. the proportion of precise vs. approximate responses was tallied for each scenario/version, and based on the results, 6 scenarios were selected that showed the greatest difference between highpr and lowpr, the precise answer being preferred in the former and the approximate answer in the latter. these were used as the basis for the main experiments. 3. experiment 1 – form and context. our first experiment investigated the influence of situational context on the social meaning of precise and imprecise numerical forms, in a partial replication and extension of bsb. the study employed the matched guise methodology common in sociolinguistics (lambert et al. 1960, campbell-kibler 2007), in which participants are exposed to linguistic stimuli in one of multiple versions or ‘guises’ and are asked to evaluate the speaker or writer on their properties or stances. 3.1. participants. a total of 371 self-reported native speakers of english aged 18-64 with u.s. ip addresses were recruited via prolific and paid £1.20 for participation. 3.2. design, materials, and procedure. the materials for the experiment were brief scenarios (selected via the pretest) in which one speaker asks a question and a second speaker answers it with a numerical expression. two factors were manipulated in a 2x2 design: context, i.e. required precision level (highpr, lowpr) and numerical form (precise, approximate). for example: lowpr: jamie has a new bicycle and is telling a friend about it. the friend is interested and wants to know more. friend: “how much did the bicycle cost? i’d love to get one like it.” highpr: jamie’s new bicycle was stolen. fortunately it was insured, so jamie has called the insurance company. insurence agent: “how much did the bicycle cost? i’ll start the paperwork right away.” jamie: “the bicycle cost $509.55 (precise) / about $500 (approximate).” the study was implemented in a 1-item, fully between subjects design (n ≈90 per condition), where each participant saw one of the six pre-tested scenarios (randomly assigned) in one of its four versions. participants completed three experimental tasks2: 2a full version of the experiment can be viewed at https://www.labvanced.com/player.html?id= 54649. proceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 373 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ task 1 participants rated the second speaker in the scenario on six dimensions of social meaning using a 7-point likert scale: competent, knowledgeable, well-prepared [competence-related], likeable, helpful [likeability-related] and pedantic. for example: based on what they say, jamie sounds: not at all competent ◦ ◦ ◦ ◦ ◦ ◦ ◦ very competent task 2 participants indicated what they inferred to be the speaker’s reason for using a precise [imprecise] numerical expression instead of a more approximate [more precise] alternative, via a free text fill-in-the-blank format. for example: you might have noticed that the speaker jamie used the expression $509.55 [about $500] rather than a less [more] precise expression such as about $500 [$509.55]. why do you think jamie answered that way? task 3 the ‘reasons for choice’ task was repeated in a multiple choice format, using fixed list of potential motivations selected based on prior research on imprecision (mühlenbernd & solt 2022): although you might have already mentioned this, which of the following, if any, do you think was a reason why jamie chose to use $509.55 [about $500] instead of a less [more] precise expression? check all that apply: • because it was the appropriate level of detail in the situation • to give as much information as possible • to avoid giving too much information • because the speaker didn’t know the exact information • to avoid saying something that might be false • to sound smart • to sound easy-going • to make the information easier to understand • because it was easier for the speaker • none of these 3.3. predictions. below we state the predictions for task 1 (social meaning ratings). tasks 2 and 3 and their correlations with the results of task 1 are discussed in section 5. based on previous findings in the literature (beltrama 2018, beltrama et al. 2022), we predict the following main effects of form: p1-1 precise will be rated higher than approximate on the competence-related attributes competent, knowledgeable and well-prepared p1-2 approximate will be rated higher than precise on the likeability-related attributes likeable and helpful p1-3 precise will be rated higher than approximate on pedantic we furthermore predict the following effects of context, i.e. interactions of form and context: proceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 374 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: ratings on social meaning attributes – experiment 1 p1-4 the positive difference between precise and approximate on competence attributes will be greater in highpr than lowpr p1-5 the positive difference between approximate and precise on likeability attributes will be greater in lowpr than highpr p1-6 the positive difference between precise and approximate on pedantic will be greater in lowpr than highpr 3.4. results. all analyses were performed using r statistical software (v4.3.2; r core team 2021). the results of task 1 are shown in figure 1. for each attribute, a linear mixed effects model was fitted to the data using the package lmertest (kuznetsova et al. 2015), with form, context and their interaction as fixed factors and random intercepts for scenario. predictors were sum coded. for each of the three competence-related attributes competent, knowledgeable and well-prepared, a significant main effect of form was found (precise higher; p<0.001 for all), supporting prediction p1-1. furthermore, there was a significant interaction of context and form (greater effect in highpr; competent/well-prepared p<0.001, knowledgeable p<0.05), supporting prediction p1-4. for the likeability-related attributes, no significant advantage for approximate over precise was found, contra prediction p1-2. rather, for likeable, there was no main effect of form, whereas for helpful, precise was rated significantly higher than approximate (p<0.001), as with the competence-related attributes. however, in both cases, there was a significant interaction of context and form (p<0.001), with the relative strength of approximate vs. precise greater in lowpr than highpr, in line with prediction p1-5. finally, for pedantic, a significant main effect of form proceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 375 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ was found (precise higher; p<0.001), supporting prediction p1-3; the interaction of form and context is not significant (p=0.16), contra p1-6, though there was a numerical difference in the predicted direction. 3.5. discussion. consistent with previous work and our predictions, we find that the choice of numerical precision level carries social meaning: precise speakers are rated higher on competencerelated dimensions than approximate ones, but are also found to be more pedantic. importantly, these effects are sensitive to contextual factors. specifically, the positive properties attributed to speakers using precise numerical forms are more pronounced in situations where high precision is called for, whereas the perceived positive associations with speaking approximately are more pronounced in low precision contexts. thus these findings support the view that the social meaning of (im)precision derives at least in part from reasoning about the needs of the context. the present study replicates and extends the findings of bsb. notably, the effects of context (i.e. interactions of form and context) found here were more consistent and robust than those observed in the previous work. we attribute this to the fact that the scenarios tested in the present experiment were selected via a pretest that established that their two contextual versions differed in required precision level as intended. the most significant difference between the present results and those of bsb is that we found no overall advantage for the approximate form on likeabilityrelated dimensions. we hypothesize that this might be due to the specific attributes included in the two studies, but also that it may reflect a difference in the structure of the experimental items. specifically, the present study, in contrast to bsb, employed a question/answer format, such that even an overly precise speaker is providing information that their interlocutor has explicitly asked for. whether such a factor indeed plays a role would need to be verified; but should this effect be found, it would be further support for the role of contextual appropriateness in deriving (especially) likeability-related associations from speakers’ choice of form, consistent with the pragmatic view advocated here. 4. experiment 2 – knowledge state. in our second experiment, we assess the effect of established speaker knowledge on the social meaning of (im)precise numerical expressions. in the scenarios tested in experiment 1, it is left unspecified whether the speaker is responding from memory or is instead consulting information available in the context. in experiment 2, we modify this by investigating how perceptions of the speaker change when it is known that he or she has a precise information source available, which has been consulted before answering. as discussed in section 1, on the pragmatic account proposed here, this manipulation is expected to amplify certain effects while diminishing others. specifically, a precise speaker who is known to be simply reading from some available information source might be seen as less competent than one reporting from memory (since high knowledge is not required) but also less pedantic (since they have a reason for speaking precisely). conversely, assuming high precision is not necessary, an approximate speaker who is known to have precise information available might be perceived as more likeable (since it can be more reliably inferred that the speaker is choosing to ‘round off’ for situational reasons rather than speaking approximately due to lack of knowledge). 4.1. participants. a total of 391 self-reported native speakers of english aged 18-64 with u.s. ip addresses were recruited via prolific and paid £1.20 for participation. proceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 376 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 4.2. materials, design, and procedure. the materials were modified versions of the same six scenarios tested in experiment 1, in which it was established that the speaker has access to a source of precise information and has consulted this before answering. as in experiment 1, two factors were manipulated – form and context – yielding four conditions in total. for example: lowpr: jamie has a new bicycle and is telling a friend about it. the friend is interested and wants to know more. jamie has pulled out the purchased receipt which has more information about the bicycle. friend: “how much did the bicycle cost? i’d love to get one like it.” highpr: jamie’s new bicycle was stolen. fortunately it was insured, so jamie has pulled out the purchase receipt and called the insurance company. insurence agent: “how much did the bicycle cost? i’ll start the paperwork right away.” jamie looks at the receipt and then answers: “the bicycle cost $509.55 (precise) / about $500 (approximate).” the experimental procedure was identical to that in experiment 1. for the analysis, data from experiment 2 were combined with those from experiment 1 to yield 8 conditions in total (2 forms x 2 contexts x 2 knowledge states). 4.3. predictions. based on the hypotheses sketched out above, we formulated four hypotheses for how the new knowledge conditions would differ from the original four base conditions: p2-1 for the three competence-related attributes, the positive difference between precise and approximate will be smaller in the knowledge condition than the base condition. p2-2 restricting to the lowpr conditions: for pedantic, the positive difference between precise and approximate will be smaller in the knowledge condition than the base condition. p2-3 restricting to the lowpr conditions: for likeable, the positive difference between approximate and precise will be greater in the knowledge condition than the base condition. p2-4 restricting to the highpr conditions: for helpful, there will be a positive difference between precise and approximate in the knowledge condition, and this will be greater than the corresponding difference in the base condition. 4.4. results. as in experiment 1, these predictions were tested by fitting a linear mixed effects model to the ratings on each attribute, with form (precise/approximate), knowledge state (knowledge/base) and their interaction as fixed effects and random intercepts for scenario. depending on the prediction, either the full data set was considered (p2-1) or the analysis was restricted to the lowpr (p2-2, p2-3) or highpr (p2-4) conditions. regarding prediction p2-1, on the competence attribute well-prepared, a significant main effect of form was again found (precise higher, p<0.001), as well as significant interaction of form and knowledge state (smaller difference in knowledge condition), p<0.05), supporting the prediction. for the attributes competent and knowledgeable, however, there was again a main effect of form (p<0.001) but no significant interaction of form and knowledge state. regarding prediction p2-2, on the attribute pedantic in the lowpr conditions, a main effect of form was again found (precise higher, p<0.001). the interaction of form and knowledge state was proceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 377 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ not significant, contra the prediction, though there was as predicted a tendency for the difference to be smaller in the knowledge condition (p=0.12). furthermore, there was a significant main effect of knowledge state, in that the speaker (regardless of form) was rated as less pedantic in the knowledge condition than the base condition; that is, established speaker knowledge had an effect across forms, and not simply for the precise speaker. prediction p2-3 was not supported. in the lowpr contexts, there was no significant effect of form (as in experiment 1) and no significant interaction of form and knowledge state. finally, prediction p2-4 was only partially supported: in the highpr context, there was a significant effect of precision in the knowledge condition, precise rated significantly more likeable approximate (p<0.001), but no significant interaction of precision and knowledge state. 4.5. discussion. in this second experiment, we sought to test the hypothesis that reasoning about a speaker’s knowledge state contributes to the derivation of social inferences, by testing scenarios in which it is established that the speaker had precise information available (for example, a purchase receipt) on which to base their response to their interlocutor’s question. we found that this modification had some effects on perceptions of the speaker, but these were less consistent and weaker than predicted. a possible conclusion from this result is that, in contradition to the predictions of our pragmatic account, reasoning about whether and how a speaker has precise knowledge of the relevant value does not in fact play a significant role in the derivation of social inferences about that speaker. there are however other possible explanations that can be considered. in particular, it might be that the experimental manipulation simply did not work as intended, that is, that participants did not perceive the speakers in the modified scenarios to have a precise information source available. alternately, the modification to the experimental scenarios might have resulted in other unexpected changes in participants’ reasoning, potentially offsetting the predicted effects. we return to this question in section 5, where we investigate participants’ perceptions of the reasons for speakers’ expression choice. 5. choice motivations. the pragmatic account of the derivation of social meaning proposed in this work holds that hearers draw inferences about the reasons for a speaker’s choice between alternative linguistic forms, and on this basis derive further social inferences about the properties of the speaker. to explore this, we elicited judgments from participants regarding the motivations they attributed to the speaker in the experimental scenario for their choice of precise or approximate form, and then investigated how these responses correlated with participants’ ratings on the six social meaning attributes. we discuss these findings here. as explained in section 2, judgments were elicited first via free text (task 2) and then via a multiple choice checkbox task (task 3). since the preregistered predictions were stated relative to the multiple choice task 3, we focus on these findings here, leaving a full discussion of the results of task 2 to future work. a summary of the results of task 3 is presented in section 5.1, and their correlation with the social meaning ratings from task 1 discussed in section 5.2. in section 5.3, we consider the implications of these findings, including how they shed light on the results of experiment 2. proceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 378 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 2: motivations 5.1. results task 3. figure 2 presents the results of task 3 for the combined data set from experiments 1 and 2, broken out by form (precise vs. approximate). as seen here, participants attributed a range of motivations to the speaker for their use of a precise or approximate numerical expression instead of its alternative. furthermore, as expected, the most commonly attributed motivations differed for the two forms. in the case of the precise form, the most frequently selected answers were: “to give as much information as possible” and “because it was the appropriate level of detail in the situation”. for the approximate form, a wider range of motivations were selected, including: “because the speaker didn’t know the exact information”, “because it was easier for the speaker”, “because it was the appropriate level of detail in the situation”, “to make the information easier to understand” and “ to avoid saying something that might be false”. two predictions were preregistered to serve as basic checks that participants’ responses on this task were sensitive to the properties of the scenario in the expected way. p3-1 context manipulation: restricting to the results of experiment 1, for the precise form, “appropriate level of detail in the situation” will be selected more frequently in the highprecision condition highpr than the low-precision condition lowpr, whereas for the the approximate form, it will be selected more frequently in lowpr than highpr. p3-2 knowledge state manipulation: for the approximate form, “speaker didn’t know the exact information” will be selected more frequently as a motivation in the original base condition than in the knowledge condition. proceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 379 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ these predictions were tested by fitting a generalized linear mixed effects model to the indicated motivation responses in the relevant data set, with context [p3-1] or knowledge state [p3-2] as a fixed factor and random intercept for scenario. both predictions were supported (p<0.001), indicating that participants were indeed able to draw context-sensitive inferences regarding the reasons behind speakers’ choices between precise and approximate numerical forms. 5.2. motivations as predictors of social meaning ratings. from the pragmatic account of social meaning, we derived and preregistered a series of predictions regarding the correlations that would be observed between the motivation(s) attributed to the speaker for their choice of form and the ratings of the speaker on the six social meaning attributes. p4-1 for the approximate form in the base knowledge state condition, the selection of “speaker didn’t know the exact information” as a motivation will correlate with lower ratings on the competence attributes competent, knowledgeable and well-prepared. p4-2 for the approximate form in the base knowledge state condition, the selection of “to make the information easier to understand” will be correlated with higher ratings on the likeability attributes likeable and helpful. p4-3 for the precise form (across contexts and knowledge states), the selection of “to sound smart” will be correlated with lower ratings on likeable. p4-4 for the precise form in the lowpr context, the selection of “to give as much information as possible” will be correlated with lower ratings on likeable and higher ratings on pedantic. in each case, the prediction was tested by fitting a linear mixed effects model to the relevant subset of the combined data from experiments 1 and 2, with ratings on the social meaning attribute from task 1 as the dependent variable, the motivation as a fixed effect and random intercept for scenario. regarding prediction p4-1, the effect of “speaker didn’t know the exact information” on the three competence attributes was not significant, contra the prediction. there was however a tendency in the predicted direction for competent and well-prepared (p=0.10 and p=0.13, respectively). prediction p4-2 was supported: the approximate speaker was rated as more likeable and helpful when “to make the information easier to understand” was selected as a reason for the choice (p<0.05 in both cases). prediction p4-3 was likewise supported: the precise speaker was rated as less likeable when “to sound smart” was selected (p<0.05). finally, prediction p4-4 was not supported: “to give as much information as possible” as a reason for use of the precise form in the lowpr context did not significantly affect ratings on likeable or pedantic. 5.3. discussion. our findings from the third experimental task provide evidence that listeners can derive inferences about the reasons for a speaker’s numerical expression choice in a given situation. furthermore, these inferred motivations were demonstrated to be predictors for how subjects evaluated those speakers with respect to their social properties. in line with our predictions, when it was inferred that the speaker used an approximate form to make the information easier to understand, this led to a perception of the speaker as more likeable, whereas when the reason was inferred to be lack of knowledge of the precise value, this yielded a tendency towards lower ratings on competence-related dimensions. similarly, when it was inferred that a speaker chose a precise form “to sound smart” (as opposed, say, in response to the needs of the situation), this lowered likeability perceptions. while the particular pattern of effects we found differed someproceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 380 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ what from the original predictions (especially for prediction p4-4), the overall picture supports the pragmatic hypothesis, according to which social meaning derives from inferences about why the speaker made a given linguistic choice. while space considerations do not allow a full discussion of the results of the free text version of the motivations task (task 2), we briefly note that these data yield a similar picture, while also identifying additional perceived motivations with predictive value. very broadly speaking, inferred reasons for choice that relate to the needs of the situation and/or the hearer correlated with more positive evaluations of the speaker, while those that related to the speaker (e.g. what is easier for the speaker or generally preferred by the speaker) led to less positive evaluations. these patterns too are consistent with the pragmatic derivation of social meaning in this case study. finally, findings from the exploration of inferred reasons for choice also shed light on the results of experiment 2, where it was seen that established speaker knowledge had less effect than predicted on social meaning ratings. one possibility raised was that the knowledge state manipulation did not work as desired. we found some support for this from the motivations task: focusing on the approximate form, “speaker didn’t know the exact information” was as predicted selected less frequently in the knowledge condition than in the base condition (per prediction p3-2); but this answer was nonetheless chosen by a sizeable minority of participants in the knowledge condition (base condition: 70%; knowledge condition: 42%). thus the scenarios tested in experiment 2 apparently did not fully rule out the possibility that the approximate speaker’s choice was in some way based on lack of precise knowledge. furthermore, the reasoning behind our predictions was that once lack of knowledge was ruled out as a motivation for the approximate choice, participants would infer that the speaker was rounding for hearer-oriented reasons. indeed,“to make the information easier to understand” was selected more frequently in the knowledge condition than the original base condition (base: 31%; knowledge: 48%); but so too was “easier for the speaker” (base: 33%; knowledge: 55%), which was not correlated with more positive ratings of the speaker. thus it seems that participants’ reasoning process did not necessarily follow the (overly) simple path we expected, and correspondingly, the effect of knowledge state on social evaluations was likewise more complex than predicted. in summary, while experiment 2 did not provide extremely strong evidence for the role of reasoning about speaker knowledge in the derivation of social inferences, neither does it provide compelling evidence against the existence of such an effect. 6. conclusions. in contrast to a common view in sociolinguistics that social meanings are indexically associated with linguistic forms themselves, our results support the view that for particular types of variation, especially those where alternatives differ in their semantic content, social meanings may derive from pragmatic inferences drawn by the hearer, including reasoning about the needs of the situation, the speaker’s epistemic state, and the reasons for their choice of form. more precisely, our findings show that the properties attributed to speakers who use precise and approximate numerical forms are sensitive to the requirements of the context of utterance (p1-4, p1-5), the inferred knowledge state of the speaker (p2-1 for well-prepared, p4-1), as well as possible reasons for the speaker’s choice of form (p4-2, p4-3). these effects all support the pragmatic account. while we have focused here on a single case study, namely that of precision level variation, we believe that the pragmatic derivation of social meaning is a more general phenomenon. in fact, beltrama & papafragou (2023) find similar effects to those discussed here for utterances in which proceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 381 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ speakers either comply with or violate the gricean maxims of relevance and informativity. in future work, we plan to investigate in greater depth the link between inferences about a speaker’s motivations for a linguistic choice and inferences about the speaker himor herself. on the basis of these findings, our further goal is to develop a formal probabilistic model that predicts social evaluations of a speaker by connecting linguistic form, context and social meaning via inferences about the speaker’s epistemic state and strategy for choosing between alternatives. references acton, eric k. 2019. pragmatics and the social life of the english definite article. language 95(1). 37–65. https://doi.org/10.1353/lan.2019.0010. beltrama, andrea. 2018. precision and speaker qualities. the social meaning of pragmatic detail. linguistics vanguard 4(1). https://doi.org/10.1515/lingvan-2018-0003. beltrama, andrea & anna papafragou. 2023. pragmatic violations affect social inferences about the speaker. glossa psycholinguistics 2(1). https://doi.org/10.5070/g601197. beltrama, andrea, stephanie solt & heather burnett. 2022. context, precision, and social perception: a sociopragmatic study. language in society 52. 805–835. https://doi.org/10.1017/s0047404522000240. burnett, heather. 2019. signalling games, sociolinguistic variation and the construction of style. linguistics and philosophy 42. 419–450. https://doi.org/10.1007/s10988-018-9254-y. campbell-kibler, kathryn. 2007. accent, (ing), and the social logic of listener perceptions. american speech 82(1). 32–64. https://doi.org/10.1215/00031283-2007-002. campbell-kibler, kathryn. 2011. the sociolinguistic variant as a carrier of social meaning. language variation and change 22. 423–441. http://dx.doi.org/10.1017/s0954394510000177. eckert, penelope. 2012. three waves of variation study: the emergence of meaning in the study of sociolinguistic variation. annual review of anthropology 41. 87–100. http://dx.doi.org/10.1146/annurev-anthro-092611-145828. van der henst, jean-baptisete, laure carles & dan sperber. 2002. truthfulness and relevance in telling the time. mind and language 81(17). 457–466. http://dx.doi.org/10.1111/14680017.00207. kuznetsova, alexandra, per bruun brockhoff & rune haubo bojesen christensen. 2015. lmertest: tests in linear mixed effects models. r package version 2.0-29. http://cran.rproject.org/package=lmertest. lambert, wallace e., richard c. hodgson, robert c. gardner & samuel fillenbaum. 1960. evaluational reactions to spoken languages. the journal of abnormal and social psychology 60(44). https://doi.org/10.1037/h0044430. mühlenbernd, roland & stephanie solt. 2022. modeling (im)precision in context. linguistics vanguard 8(1). 113–127. https://doi.org/10.1515/lingvan-2022-0035. r core team. 2021. r: a language and environment for statistical computing. r foundation for statistical computing, vienna, austria. http://www.r-project.org/. solt, stephanie, chris cummins & marijan palmović. 2017. the preference for approximation. international review of pragmatics 9(2). 248–268. https://doi.org/10.1163/1877310900901010. proceedings of elm 3: 371-382, 2025 stephanie solt, roland mühlenbernd, and mariya burbelko: social meaning and pragmatic reasoning: the case of (im)precision. 382 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ a conceptual analysis of verbs of pushing and pulling anton benz, torgrim solstad, oliver bott, martin kahnberg, andrea c. schalley* abstract. although verbal expressions of caused motion, such as push and pull, have been extensively studied within linguistics, semantic dimensions beyond path and manner of motion have received less attention. this pilot study aims to identify such dimensions involved in the expression of caused motion in german, focusing on observable properties of pushing and pulling events that determine the selection of verbs to describe events of caused motion. using 3d graphical modeling, participants were presented with video clips of a computer-animated agent moving a barrel, thereby allowing for a systematic manipulation of properties and hence dimensions. we investigated four dimensions to assess their impact on verb selection: (i) angle of contact, (ii) movement of the agent relative to the barrel, (iii) the agent’s orientation/facing, and (iv) the force employed. cluster and principal component analyses were conducted on the collected linguistic data. verbs were represented by five-dimensional vectors capturing correlations with the cosine and sine of the angle, and marginal probabilities in conditions of instantaneous movement, forward facing, and heavy force. our findings indicate that conceptually distinguishable verb clusters are primarily defined by the movement feature – that is, whether the agent moves together with the barrel or not – and the cosine of the angle. contrary to theoretical predictions, little evidence was found supporting the categorization of verbs based on the force applied to the barrel. these results suggest that the movement and position of the agent relative to the moved object are key determinants in the production of verbal descriptions of caused motion events. keywords. verbs of pushing and pulling; caused motion; conceptualization; lexicalization; german 1. introduction. dating back to seminal work by talmy (cf. e.g., talmy 1985), verbal expressions of caused motion (e.g., put, push) have been studied extensively within linguistics from the perspective of linguistic typology, theoretical linguistics, language acquisition and processing (cf. e.g., allen et al. 2007, goldberg 1995, hendriks et al. 2008, margetts et al. 2022). whereas research in the tradition of talmy has investigated the realization of motion and path components in terms of sateliteand verb-framed languages, another important line of research has focused on the argument-structural realization and alternations associated with these verbs (e.g., levin 1993, goldberg 1995; and much subsequent work). however, less attention has been directed towards semantic dimensions beyond path and manner of motion at play in these verbs. thus, while push and pull may both be considered to be *special thanks to our student assistants sophia höbel who annotated the data and to henry salfner for assisting us in the creation of the animations. torgrim solstad and oliver bott gratefully acknowledge funding by the deutsche forschungsgemeinschaft (dfg, german research foundation) to subproject b01, crc 1646, project number 512393437. authors: anton benz, leibniz-centre general linguistics, zas (benz@leibniz-zas.de) & torgrim solstad, bielefeld university (torgrim.solstad@uni-bielefeld.de) & oliver bott, bielefeld university (oliver.bott@unibielefeld.de) & martin kahnberg, karlstad university (kahnberg@gmail.com) & andrea c. schalley, karlstad university (andrea.schalley@kau.se). proceedings of elm 3: 43-52, 2025 c©2025 anton benz, torgrim solstad, oliver bott, martin kahnberg, and andrea c. schalley published by the lsa with permission of the author(s) under a cc by license. 43 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ verbs of caused motion, they differ in the relation of the agent and the moved object. for pull, a force is exerted in the direction of the agent, whereas for push, the force is directed away from the agent. what is more, some verbs seem to encode a stronger application of force than others, as witnessed by the difference between pull, drag and haul. in this paper, we present a pilot study using 3d graphical modelling to elicit linguistic productions of relevance to some of the dimensions involved. this pilot study is part of a larger endeavour in which we aim to identify more precisely the semantic dimensions involved in the expression of caused motion across languages. more specifically, we ask which observable properties of pushing and pulling events determine which verb is best used to describe them, and how they can be represented. what are the observable properties of agent, moved object and their spatial relation? one goal of our investigation is to model the semantic dimensions of caused motion verbs as geometric structures within the conceputal spaces framework (gärdenfors 2000). in gärdenfors’ theory, concepts are analysed as regions in multi-dimensional spaces which are derived from (fine-grained) semantic dimensions that to a large extent are based on perception (gärdenfors 2000, gärdenfors & warglien 2012). the geometric structures provide an objective, languageindependent measure of semantic similarity. previous research has provided profound evidence for a geometrical organization of concepts in the (direct) sensory domain, such as colour (roberson et al. 2005), olfaction (majid et al. 2018), static spatial relations (levinson & wilkins 2006), and even prototypical instances of motion events (giese et al. 2008, malt et al. 2014). recent work (gärdenfors 2020, gärdenfors et al. 2018, warglien et al. 2012, wolff 2007) extended gärdenfors’ approach in the verbal domain and provided ideas for integrating structural accounts. however, less progress has been made in the conceptual space of caused motion events involving both an agent and a patient as for verbs like push and pull. past studies in the domain of caused motion have used 2d videos to elicit descriptions of basic pushing and pulling events, focusing on the difference between verband satellite-framed languages (e.g. hickmann et al. 2018, montero-melis 2021). based on a production experiment using short 3d video clips, the present study instead aimed at assessing in more detail which (observable) semantic dimensions make out the domain of pushing and pulling, which we see as a fundamental domain of physical interaction between agents and patients. furthermore, previous studies focused on prototypical events, using human actors (e.g. malt et al. 2014) or idealized animations (e.g. montero-melis 2021). however, one needs peripheral event instances in order to pinpoint conceptual boundaries. a systematic manipulation of several dimensions that moves from prototypical to peripheral event instances leads to a large number of combinatorial possibilities to be tested. we addressed this challenge by presenting participants with video clips in which a computer-animated agent moved a barrel over a short distance, thereby allowing for fine-grained adjustments of potentially impactful properties. it should be noted that the main research goal in this pilot study was to determine the predictors that trigger the production of different verbs and to classify the verbs in semantic verb clusters. the role of modifiers, or — in talmy’s terms — satellites, of various types is not discussed in this paper, but is part of a more detailed analysis that is still on-going. 2. the experiment. our experiment investigates how physical properties of pushing and pulling events influence the speaker’s choice of verb (and modifiers). as mentioned before, force is a sigproceedings of elm 3: 43-52, 2025 anton benz, torgrim solstad, oliver bott, martin kahnberg, and andrea c. schalley: a conceptual analysis of verbs of pushing and pulling. 44 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ nificant concept in cognitive linguistics, particularly in talmy’s work. however, it is a multifaceted concept rooted in physics. in physics, force is a quantity that has both magnitude and direction and is represented by a vector. in caused motion, force is exerted over a period of time along a part or the entire length of the path taken by the moving object. therefore, when conceptualizing force in a caused motion event, we can consider the strength (magnitude) of the force that an agent applies to an object, and the length of time during which this force is applied—is it instantaneous or continuous over the entire path length? for the perception of force, the physical orientation of the agent relative to the object is also a relevant aspect that can determine how the event is verbalized. in terms of physical orientation, we distinguish between the angle between the agent and the object, and whether the agent’s body is oriented towards the object or in the direction of movement. we, therefore, identified four dimensions to test their impact on verb selection. the first dimension concerns the angle of contact between the agent and the barrel’s direction of movement. we expected verb meanings to be sensitive to the agent’s relative position to the barrel. in prototypical pushing events, the agent is positioned directly behind the moving object (0◦ angle). in prototypical pulling events, on the other hand, the agent is directly in front of the object (180◦ angle). to obtain a fuller picture of verb choice, we added intermediary angles in 45◦ intervals. based on native speaker intuitions, we expected speakers to switch from push to pull type verbs in the range of 90◦ to 135◦. for this reason, we divided this range more finely into 15◦ intervals. consequently, our study included seven different angles: 0◦, 45◦, 90◦, 105◦, 120◦, 135◦, and 180◦. figure 1 illustrates four of these angles (180◦, 105◦, 135◦ and 0◦). examples of selected videos illustrating this and the other dimensions discussed below can be found in the associated osf archive.1 figure 1: stills for 180◦/105◦/135◦/0◦ with left/right movement and agent facing forwards/to object the second dimension involves the movement of the agent, who can either remain in place, that is, only move their arms – resulting in instantaneous contact with the object – or move along with the object, leading to continuous contact. for example, a prototypical instance of schubsen ‘shove’ would be instantaneous, while gehen mit ‘walk (along) with’ would involve continuous movement. the third dimension we manipulated is force, which can be light or heavy. we expected this dimension to distinguish between uses of the german push type verbs, with schubsen ‘push, shove’ being light in strength and stoßen ‘thrust, shove’ being heavy. we characterized the fourth dimension as the agent’s facing. we manipulated two different agent postures: one where the agent faces towards the object, and one where the agent faces in the direction of the barrel’s movement. the intuition behind this was that the body posture of the agent would reflect their attention and therefore be relevant for the perception of the agent’s effort. 1https://osf.io/mfstn/?view only=4d5736ce36c547df94278e8d5188ac28 proceedings of elm 3: 43-52, 2025 anton benz, torgrim solstad, oliver bott, martin kahnberg, and andrea c. schalley: a conceptual analysis of verbs of pushing and pulling. 45 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ finally, as a counter-balancing factor, we also created two version of all videos, one in which the direction of the barrel’s movement across the screen was manipulated, either to the right or to the left on the screen, as seen from participant’s perspective. it should be noted that we chose to use an inanimate object as the moved object to avoid any complications with regard to the interaction between agent and patient. ultimately, however, it would be desirable to include objects that have varying properties of ‘self-propelling’, such as objects on wheels or even another animate entity. 2.1. methods. 2.1.1. participants. we recruited 81 german native speakers (36 male, 45 female; mean age: 24.5 years) via the platform prolific. participants were directed to our server where the experiment was hosted. 2.1.2. design. based on the above dimensions, the experiment employed a 7×2×2×2 withinparticipants and within-items factorial design manipulating the following factors: angle between the agent and the barrel’s direction of movement (seven angles: 0◦, 45◦, 90◦, 105◦, 120◦, 135◦, 180◦), movement type (continuous vs. instantaneous), facing (towards barrel vs. forwards in direction of movement), force applied (light vs. heavy). in addition, direction of the barrel’s movement was included as a counter-balancing factor (to the right vs. to the left on the screen). this design yields 56 distinct main conditions (excluding direction). however, since there is no physical distinction between facing forward and facing the object for the 0◦ angle, we obtained a total of 52 visually discernible main conditions. the dependent variable reported in this paper was the verb produced. 2.1.3. stimuli. including the left vs. right manipulation involved for direction, we created 104 animated video clips, each approximately 3 seconds long, depicting a human-like agent pushing or pulling a barrel (see tab. 1 and figure 1). the videos were generated using a state-of-the-art physics engine (blender, https://www.blender.org/). the animations were distributed on two lists such that all 52 (distinguishable) conditions involving angle, movement, force and facing were in both lists. half of the left or right direction conditions were included in one list, the rest in the other. the order of conditions was pseudo-randomized, seeking to minimize repetition across factors. for example, two subsequent trials always differed by at least 45◦ in angle to make videos more easily discernible. finally, a reversed order version of each of these lists was created, bringing the number of experimental lists to four. 2.1.4. procedure. the experiment was implemented using the free onexp software (version 1.3.1).2 after reading short written instructions, participants proceeded to a short practice of two trials, upon which they received the experiment in 52 trials in a single block. participants watched video clips of approximately 3 seconds each, presented individually on separate pages, showing an agent pushing or pulling a barrel. the videos were displayed centrally on their screen, accompanied by the question “what does the person do with the barrel?”. below this, a sentence frame with a text field was provided for them to complete, see (1). their task was to describe as spontaneously as possible what the agent was doing. before starting the main experiment, participants 2see http://onexp.textstrukturen.uni-goettingen.de proceedings of elm 3: 43-52, 2025 anton benz, torgrim solstad, oliver bott, martin kahnberg, and andrea c. schalley: a conceptual analysis of verbs of pushing and pulling. 46 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ were informed that their descriptions would later be used to characterize the individual video clips. therefore, it would be important that their descriptions be specific enough for another person to assign them to the correct video clips. they were also instructed that each video clip was independent, featuring a different animated person, and that the barrel might behave differently in each clip. (1) the person (original in german) 2.2. data and coding. we gathered a corpus of 4,212 descriptions (word length range: 3–70, mean 8.7). we annotated >20 features of the descriptions (not yet finalised). in the following analysis, we concentrated on 3 features: (1) the main verbs in the matrix clauses that express movement of the barrel, (2) whether the main verb is a particle verb or not, and (3) whether the barrel (the moved object) is the direct object or embedded in a prepositional phrase (pp). we found 95 different matrix verb constructions with 9 matrix verbs that have a frequency > 0.5%, see (2).3 (2) ziehen ‘pull’ 1635 schieben ‘push’ 1156 drücken ‘press’ 195 schubsen ’push’, ‘shove’ 195 stoßen ‘thrust, shove’ 176 gehen ‘walk’ 173 bewegen (refl) ‘move (oneself)’ 102 bewegen ‘move’ 71 laufen ‘walk’ 29 2.3. descriptive results. based on introspection, we identified four types of event descriptions, see figure 2. the first two types conceptualize the movement event as direct causation involving a force exerted on the barrel. they are distinguished by the direction of the force vector – either towards or away from the agent. the third type describes the agent’s movement, with the barrel following as a satellite. the fourth type describes the movements of the agent and the barrel in two separate clauses, conceptualizing their respective movements as two parallel events. examples for all types are shown in (3) (originals in german). (3) the person . . . a. simply pushes the barrel to the left. b. pulls the barrel behind him with one arm. c. walks with the barrel to the right. d. walks to the left and pushes the barrel next to him. inspection of the frequency data (see table 1) reveals further regularities. verbs of the ‘walkwith’ type occur only with continuous movement and are most frequent at intermediate angles. the ‘shove’ type verbs schubsen and schieben occur only with instantaneous movements and are most frequent at lower angles. other verbs show a steady increase or decrease in frequency with angle and can occur with both instantaneous and continuous movement. 2.4. analytic results. for the cluster analysis, each verb was represented by a 5-dimensional vector, capturing its correlation with the cosine and sine of the angle, and its marginal probabili3see the osf archive https://osf.io/mfstn/?view only=4d5736ce36c547df94278e8d5188ac28 for a list with data. proceedings of elm 3: 43-52, 2025 anton benz, torgrim solstad, oliver bott, martin kahnberg, and andrea c. schalley: a conceptual analysis of verbs of pushing and pulling. 47 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ a. schubsen, schieben b. ziehen shove, push pull force move force move c. laufen mit d. gehen & schieben walk with walk & push move causes movement with move move force figure 2: classification of event descriptions ties of being produced in three conditions: instantaneous movement (meanmov), forward facing (meanfacing), and heavy force (meanforce). all features were coded as binary variables (1 indicating presence). a positive correlation with cosine indicates that verb frequency decreases as the angle increases, whereas a negative correlation signifies an increase in frequency as the angle decreases. a positive correlation with sine suggests that the verb is most frequently used for angles near 90◦. a principal component analysis (pca) was conducted to examine the underlying structure of the dataset. according to kaiser’s criterion (eigenvalue > 1), the optimal dimensionality would be 2 components. however, adopting a more lenient threshold (eigenvalue > 0.7), 3 dimensions are justified. we chose to retain 3 principal components, which together explain 91% of the total variance (with the first 2 components explaining 73%). the factor loading plot and a score plot of individual verbs for the first two components are presented in figure 3. the factor loading plot indicates that the first component is primarily influenced by correlations with sine, cosine, and meanmov. additionally, a heatmap depicting clusters based on these features shows that verbs are predominantly grouped according to their correlation with these three variables, see figure 4. figure 3: factor loading plot and score plot for dimensions 1 & 2 of the pca analysis. k-means clustering (k = 3) was performed for binary feature combinations (see figure 4). the movement feature (continuous vs. instantaneous) identified three verb clusters: verbs such as bewegen (refl.) ‘move’, gehen ‘walk’, and laufen ‘walk’ were associated with the +continuous proceedings of elm 3: 43-52, 2025 anton benz, torgrim solstad, oliver bott, martin kahnberg, and andrea c. schalley: a conceptual analysis of verbs of pushing and pulling. 48 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ move with walk with walk with move pull push shove shove press pull toward continuous movement instantaneous movement table 1: production probabilities of frequent verbs conditional on movement features. probabilities p (verb|angle): proportion of descriptions with matrix verb given the angle (0◦, 45◦, 90◦, 105◦, 120◦, 135◦, 180◦). no line or grey dot: angle not produced for angle and movement. coding of verb features: ‘verb-0.-0.’: no particle verb, barrel direct object; ‘verb-0.-pp’: no particle verb, barrel realized in a pp feature, while schubsen ‘push’ and stoßen ‘thrust’ were marked as +instantaneous. the remaining verbs were unmarked with respect to the continuous/instantaneous distinction (figure 4c). for all −continuous verbs, the barrel was realized as the direct object, whereas for all +continuous verbs, it was embedded in a prepositional phrase (e.g., ‘move oneself with the barrel’). a corresponding cluster analysis based on sine and cosine showed that the +continuous verbs formed a cluster positively correlated with sine, while the +instantaneous verbs formed a cluster positively correlated with cosine. verbs unmarked with respect to meanmov, with the exception of ziehen ‘pull’, were grouped in the +cosine cluster. ziehen was isolated and formed its own −cosine cluster. the features facing and force did not result in clearly delineated clusters. the results of the cluster analyses are summarised in table 2. the cluster analysis confirmed the conceptual classification of event descriptions based on introspection and semantic intuition shown in table 2. it added a finer distinction in the class of verbs expressing direct causation with a force applied in the direction of movement by differentiating between verbs that only occur in +instantaneous contexts, and verbs that are unmarked with proceedings of elm 3: 43-52, 2025 anton benz, torgrim solstad, oliver bott, martin kahnberg, and andrea c. schalley: a conceptual analysis of verbs of pushing and pulling. 49 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ x y r cos sin θ movement a) b) c) d) figure 4: cluster analyses: (a) heatmap showing clustering of verbs based on their correlations with sine, cosine, meanmov, meanforce, and meanface. (b) – (d) k-means clustering results (k = 3): (b) clusters based on correlations with sine and cosine only; (c) clusters based on meanmov and meanforce values; (d) clusters based on meanforce and meanface values. clear separation of clusters is observed in (b) and (c), with no separation visible in (d) respect to movement. for verbs produced by at least 15 participants, we fitted linear mixed effect models with cos angle (cos), movement,4 force, and facing as fixed effects and by-participant random intercepts. predictors vary for individual verbs. we found the following patterns for cos: cos was no significant predictor for +continuous-verbs; all other verbs were either positively or negatively correlated with cos (table 2), except bewegen ‘move’, which also did not correlate with cos. the other predictors may correlate with individual verbs, but we found no general pattern correlated to verb clusters. 3. discussion. in this paper, we presented the results of a pilot study in which we elicited verbal descriptions for 3d video animations of caused motion events involving an animate agent and an inanimate patient (a barrel). we manipulated four factors in the animations, (i) the angle between 4we dropped movement for +contand +inst-verbs. to resolve convergence issues, angle was transformed using the cosine function. the sine of the angle was excluded from the list of fixed effects because the unequal number of conditions at lower and higher angles led to spurious significant results in the linear model due to data asymmetry. proceedings of elm 3: 43-52, 2025 anton benz, torgrim solstad, oliver bott, martin kahnberg, and andrea c. schalley: a conceptual analysis of verbs of pushing and pulling. 50 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ move causes movement with causation move instantaneous move + continuous unmarked + instantaneous + sin angle + cos angle − cos angle + cos angle bewegen (refl)-pp (move self with) schieben (push) ziehen (pull) stoßen (push) gehen-pp (walk with) drücken (press) schubsen (push) table 2: verb clusters: semantic feature (red), use correlated (blue) agent and patient, (ii) the continuity of the agent’s movement, (iii) the force employed, (iv) and the agent’s facing, that is, the direction of their attention. for these factors, conceptually clearly distinguishable verb clusters can only be defined by the movement (continuity) feature, which tells us whether the agent moves together with the barrel (+cont) or remains in place (+inst), and the cos of the angle. interestingly, the results provide little evidence that verbs are categorized according to the force applied to the barrel (as predicted by gärdenfors & warglien’s 2012 theory). it is rather the movement and position of the agent in relation to the barrel that determine production of verbal descriptions. our ongoing coding efforts will allow us to assess the influence of the manipulated factors on a number of modifiers, including – but not limited to – the force of the agent, the source, path, and goal of the movement, or its manner. going beyond german, by testing english, italian, russian, swedish and persian, we will also develop a basis for a more broad typological investigation of the semantic dimensions of caused motion. references allen, shanley, aslı özyürek, sotaro kita, amanda brown, reyhan furman, tomoko ishizuka & mihoko fujii. 2007. language-specific and universal influences in children’s syntactic packaging of manner and path: a comparison of english, japanese, and turkish. cognition 102(1). 16–48. https://doi.org/10.1016/j.cognition.2005.12.006. giese, martin a., ian thornton & shimon edelman. 2008. metrics of the perception of body movement. journal of vision 8(9). 13–13. https://doi.org/10.1167/8.9.13. goldberg, adele e. 1995. constructions: a construction grammar approach to argument structure. chicago: university of chicago press. gärdenfors, peter. 2000. conceptual spaces: the geometry of thought. cambridge, ma: the mit press. gärdenfors, peter. 2020. events and causal mappings modeled in conceptual spaces. frontiers in psychology 11. https://doi.org/10.3389/fpsyg.2020.00630. gärdenfors, peter, jürgen jost & massimo warglien. 2018. from actions to effects: three constraints on event mappings. frontiers in psychology 9. https://doi.org/10.3389/fpsyg.2018.01391. proceedings of elm 3: 43-52, 2025 anton benz, torgrim solstad, oliver bott, martin kahnberg, and andrea c. schalley: a conceptual analysis of verbs of pushing and pulling. 51 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ gärdenfors, peter & massimo warglien. 2012. using conceptual spaces to model actions and events. journal of semantics 29(4). 487–519. https://doi.org/10.1093/jos/ffs007. hendriks, henriette, maya hickmann & annie-claude demagny. 2008. how adult english learners of french express caused motion: a comparison with english and french natives. acquisition et interaction en langue étrangère 27(15-41). hickmann, maya, henriëtte hendriks, anne-katharina harr & philippe bonnet. 2018. caused motion across child languages: a comparison of english, german, and french. journal of child language 45(6). 1247–1274. levin, beth. 1993. english verb classes and alternations: a preliminary investigation. chicago: university of chicago press. levinson, stephen c. & david p. wilkins (eds.). 2006. grammars of space: explorations in cognitive diversity. cambridge: cambridge university press. majid, a., s. g. roberts, l. cilissen, k. emmorey, b. nicodemus, l. o’grady, b. woll, b. lelan, h. de sousa, b. l. cansler, s. shayan, c. de vos, g. senft, n. j. enfield, r. a. razak, s. fedden, s. tufvesson, m. dingemanse, ö. öztürk, p. brown, c. hill, o. le guen, v. hirtzel, r. van gijn, m. a. sicoli & s. c. levinson. 2018. differential coding of perception in the world’s languages. proceedings of the national academy of sciences u.s.a. 115(45). 11369– 11376. malt, barbara c., eef ameel, mutsumi imai, silvia p. gennari, noburo saji & asifa majid. 2014. human locomotion in languages: constraints on moving and meaning. journal of memory and language 74. 107–123. margetts, anna, sonja riesberg & birgit hellwig (eds.). 2022. caused accompanied motion: bringing and taking events in a cross-linguistic perspective typological studies in language. amsterdam / philadelphia: john benjamins. https://doi.org/10.1075/tsl.134. montero-melis, guillermo. 2021. consistency in motion event encoding across languages. frontiers in psychology 12(625153). roberson, d., j. davidoff, i.r.l. davies & l.r. shapiro. 2005. color categories: evidence for the cultural relativity hypothesis. cognitive psychology 50(4). 378–411. talmy, leonard. 1985. lexicalization patterns: semantic structure in lexical forms. in timothy shopen (ed.), language typology and syntactic description iii: grammatical categories and the lexicon, 57–149. cambridge: cambridge university press. warglien, massimo, peter gärdenfors & matthijs westera. 2012. event structure, conceptual spaces and the semantics of verbs. theoretical linguistics 159–193. wolff, phillip. 2007. representing causation. journal of experimental psychology: general 136(1). 82–111. proceedings of elm 3: 43-52, 2025 anton benz, torgrim solstad, oliver bott, martin kahnberg, and andrea c. schalley: a conceptual analysis of verbs of pushing and pulling. 52 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the role of definiteness in ad hoc implicatures andré eliatamby & lyn tieu* abstract. this study investigates how ad-hoc implicatures and the definiteness presupposition of the definite determiner ‘the’ interact. using a truth value judgment task (crain & thornton 2000), we examine whether english-speaking adults interpret the definite and indefinite determiner differently in sentence pairs such as: ‘mary bought a striped sweater’ and ‘mary bought the striped sweater’, in contexts in which there are two possible referents, one which is best described with one adjective (e.g. ‘striped’) and the other which is best described with two adjectives (e.g. ‘striped’ and ‘spotted’). we find more ad hoc implicature for ‘the’ than ‘a’; that is, uses of the definite ‘the’ are rejected more frequently than uses of ‘a’ when the purchased item would best be described with two adjectives. we take this finding to suggest that the need to satisfy the uniqueness presupposition of ‘the’ acts as an additional trigger for implicature generation. this result raises questions for both neo-gricean and localist models of implicature generation, which we briefly outline. keywords. definiteness; indefinites; ad hoc implicature; presupposition; uniqueness; semantics; pragmatics; psycholinguistics 1. introduction. this paper is interested in how the uniqueness property of english articles influences the generation of ad-hoc implicatures (hirschberg 1991). it is well-known that in english, the definite and indefinite articles differ from each other with respect to uniqueness. take the sentences in (1) and (2). (1) mary bought the striped sweater. (2) mary bought a striped sweater. sentence (1) communicates that there is a unique striped sweater in the context and is therefore infelicitous in contexts in which there are two striped sweaters. sentence (2), on the other hand, can be used felicitously in contexts where there are two striped sweaters. following heim (1991), we will treat the uniqueness property of ‘the’ as a presupposition. consider now the use of (1) and (2) in the following contexts: (3) the store has a plain sweater, a sweater with stripes and spots, and a sweater that only has stripes. mary buys one sweater, namely: a. the striped and spotted sweater. b. the sweater with only stripes. do (1) and (2) have equivalent meanings in the contexts in (3a) and (3b)? our intuition suggests that (1) is only acceptable in (3b), while (2) is acceptable in both (3a) and (3b). this asymmetry is unexpected if both (1) and (2) generate ad-hoc implicatures at the matrix level, since in both constructions, ad-hoc implicatures should implicate that mary bought the sweater with no stripes. in * authors: andré eliatamby, rutgers university (ae644@ruccs.rutgers.edu) & lyn tieu, university of toronto (lyn.tieu@utoronto.ca). proceedings of elm 3: 144-151, 2025 c©2025 andré eliatamby and lyn tieu published by the lsa with permission of the author(s) under a cc by license. 144 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the above context, (1) and (2) have (4) and (5) as alternatives, which when treated as false yield (6) and (7): (4) mary bought the striped and spotted sweater. (5) mary bought a striped and spotted sweater. (6) mary didn’t buy the striped and spotted sweater. (7) mary didn’t buy a striped and spotted sweater. conjoining (6) with (1) and (7) with (2) yields the inference that the sweater mary bought had stripes and did not have spots. if our intuition is attested, something must interfere with implicature computation that leads to fewer implicatures for (2) than for (1). since the only difference between (1) and (2) is the choice of determiner, the properties of the determiner must play a role. in this paper, we test this intuition using two truth value judgement tasks with englishspeaking adults. 2. experiment 1 experiment 1 investigates how adult english speakers assess sentences like (1) and (2) in contexts where someone buys an item that is best described using two properties. 2.1. method. 2.1.1. participants. we recruited 60 english native speakers through prolific (https://www.prolific.com/) and randomly assigned them to either the ‘a’ or ‘the’ condition. all participants self-identified as native speakers of english, with normal or corrected-to-normal vision. participants were paid £1 for the task, which took on average 5m36s to complete. 2.1.2. procedure. the task was a truth value judgment task (crain & thornton 2000), implemented and hosted on qualtrics. participants were given a back story about characters who were shopping at the store. on each trial, they saw a picture containing three items, and a shopping basket under one of the items. a puppet named raffie described which item the character had purchased (using either a definite or an indefinite description), and participants had to indicate whether raffie was right or wrong by clicking on ‘yes’ or ‘no’ (see figure 1). 2.1.3 materials. noun phrase type (definite ‘the’ vs. indefinite ‘a’) was a between-subject variable. critical target trials involved weak/under-informative descriptions containing one adjective, such as ‘mary bought {a/the} striped sweater’ to describe a context in which there was both a sweater with stripes and a sweater with stripes and spots, and mary had bought the one with stripes and spots (see figure 1). if participants computed the ad-hoc implicature that the sweater mary bought didn’t contain spots, they were expected to reject the test sentence; if not, they would accept the test sentence on its literal meaning. the experiment also included unambiguously true and unambiguously false 1and 2-adjective controls, in which the test sentences were clearly true or clearly false descriptions of the purchased item (see figures 2 and 3 for examples). we also included clearly true and clearly false filler items which involved descriptions that did not contain any adjectives (e.g., ‘tara bought the carrot’). in all, each participant saw 2 training items, followed by 30 test items: 12 ambiguous target trials containing either ‘a’ or ‘the’, 6 clearly true/clearly false 1-adjective controls, 6 clearly true or clearly false 2-adjective controls, and 6 adjective-less fillers. the 2-adjective controls were presented in a second block, so as not to proceedings of elm 3: 144-151, 2025 andré eliatamby and lyn tieu: the role of definiteness in ad hoc implicatures. 145 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ interfere with participants’ interpretation of the 1-adjective targets (for instance, hearing a 2-adjective description might bias people towards rejecting the 1-adjective targets for being underinformative). within each block, all trials were presented in a completely randomized order. figure 1: screen capture of a critical target trial in the ‘the’ condition. in the ‘a’ condition, the sentence contained the indefinite determiner ‘a’ instead of ‘the’. figure 2: image from a false 1-adjective control trial, paired with the sentence ‘bernard bought {a/the} polka-dotted bag.’ proceedings of elm 3: 144-151, 2025 andré eliatamby and lyn tieu: the role of definiteness in ad hoc implicatures. 146 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 3: image from a true 2-adjective control trial, paired with the sentence: ‘ellie bought {a/the} rainbow-coloured and polka-dotted dress.’ 2.2. results. one participant was excluded for failing to score at least 12/18 (two thirds) accuracy on the unambiguous control and filler trials, leaving a total of 59 participants for analysis (29 in the ‘a’ condition and 30 in the ‘the’ condition). for these participants, accuracy was above 92% for all unambiguous filler and control conditions. figure 4 displays the average proportion of yesresponses in the target ‘a’ and ‘the’ conditions (dots represent individual participant means). mean acceptance in the indefinite ‘a’ condition was 93%, compared with 56% in the definite ‘the’ condition. we fit a mixed effect logistic regression model on responses to the target conditions, with definiteness as a fixed effect, and random intercepts for subject and item. model comparisons revealed a significant effect of definiteness (χ2(1)=16, p<.001), with participants more likely to reject the underinformative target statements when they contained the definite article than when they contained the indefinite article. figure 4: mean acceptance of critical (underinformative) 1-adjective target trials for ‘a’ and ‘the’ conditions in experiment 1. 2.3. discussion. the results of experiment 1 show that in contexts like (3a), adults were more likely to reject sentences like ‘mary bought the striped sweater’ than ‘mary bought a striped sweater’, suggesting that ad-hoc implicatures were more likely to be computed for sentences proceedings of elm 3: 144-151, 2025 andré eliatamby and lyn tieu: the role of definiteness in ad hoc implicatures. 147 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ containing the definite determiner compared to sentences containing the indefinite determiner. in fact, sentences with ‘a’ seem to rarely generate ad hoc implicatures: the majority of participants always accepted the critical ‘a’ sentences, and no participant rejected the critical indefinite sentences more than half the time. sentences like ‘mary bought a striped sweater’ seem to be generally acceptable when the sweater has both stripes and spots. while we find evidence of an asymmetry, the rate of implicature generation in the definite condition is lower than our linguistic intuition predicts; adults as a group are only at chance to reject the critical sentences. we were concerned this might be due to the stimuli being presented in a written/textual form, rather than spoken. prosodic structure is known to influence the rate at which implicatures are generated and influences the alternative set used in implicature computation (tomlinson et al. 2017, tomlinson & ronderos 2021). for example, both (8) and (9) negate different alternatives, which are compatible with mary buying a striped and spotted sweater: (8) mary bought the striped sweaterf. ↝ mary didn’t buy the striped box, the striped bag, the striped umbrella, etc. (9) mary bought the stripedf sweater. ↝ mary didn’t buy the plain sweater. without an explicit prosodic structure, participants might have parsed the critical sentences as in (8) or (9), or participants might have been unsure about how to parse the critical sentences, and as a result failed to strengthen their meanings. our design also contains a possible confound regarding the rejection of the critical sentences in the definite condition. in a context where there is a plain sweater, a sweater with stripes and spots, and a sweater that only has stripes, the unenriched denotation of ‘the striped sweater’ is undefined, since the set of striped sweaters is not a singleton set. since participants could only give true/false (yes/no) responses, rejections of the critical targets might have been due to presupposition failure, rather than participants’ belief that the enriched interpretation was falsified. since ‘a’ does not trigger a uniqueness presupposition, the asymmetry between ‘a’ and ‘the’ could simply be due to the uniqueness presupposition carried by ‘the’. if rejections of the critical trials were due to presupposition failure, then participants should reject sentences like (1) in contexts in which mary bought a striped sweater with no spots. in this case, an enriched interpretation of (1) (with an ad hoc implicature) is true, but the unenriched interpretation is still undefined. rejection of such trials would be a clear indication that presupposition failure was driving the rejection of the critical trials in our definite condition in experiment 1. in the design of experiment 1, each participant did see one (but only one) such trial. in aggregate, participants accepted these trials 100% of the time. in experiment 2, we decided to address both the absence of explicit prosody and increase the number of trials testing for presupposition failure. 3. experiment 2. in experiment 2, we addressed the absence of explicit prosody and the lack of sufficient trials testing for presupposition failure. rather than being presented as text, sentences were presented in pre-recorded videos, produced by a talking rabbit puppet. we also changed the distribution of trial types. the number of trials where the sentence mentioned one property, but the purchased item had two properties, was reduced from 12 to 6. the number of trials where the sentence mentioned one property, and the purchased item had that property, was increased from 1 to 6. proceedings of elm 3: 144-151, 2025 andré eliatamby and lyn tieu: the role of definiteness in ad hoc implicatures. 148 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3.1. method. 3.1.1. participants. 61 participants (who had not completed experiment 1) were recruited through prolific; 31 were randomly assigned to the ‘a’ condition and 30 to the ‘the’ condition. all self-reported as native speakers of english, with normal or corrected-to-normal vision. participants were paid £1.35 for the task, which took on average 8m47s to complete. 3.1.2 procedure. the procedure was the same as in experiment 1, with participants judging whether a puppet’s descriptions of the pictured scenarios were true or false. there were two differences between the two experiments. first, test sentences in experiment 2 were presented orally (the puppet produced the sentences in pre-recorded videos). second, we decreased the number of critical targets to add 1-adjective true controls (see figure 5), which were presented in a second block following the critical trials. figure 5: example image associated with a 1-adjective true trial in experiment 2. this image accompanied the sentence: ‘evan bought {a/the} flowery pillow.’ 3.1.3 materials. the materials for experiment 2 were very similar to those in experiment 1, except for the changes mentioned above. in total, participants saw two training trials, followed by 30 test trials presented in two test blocks: the first block contained a completely randomized sequence of six 1-adjective-true critical targets, 6 clearly true or clearly false controls, and 6 true/false fillers; the second block contained six 1-adjective true targets and six 2-adjective true/false controls. 3.2. results. 53 participants scored at least 2/3 accuracy on the unambiguous controls and fillers and were retained for analysis (27 ‘a’, 26 ‘the’). for these participants, accuracy on the unambiguous controls and fillers was above 92%. the left-hand side of figure 6 displays the average proportion of yes-responses in the target ‘a’ and ‘the’ conditions (dots represent individual participant means). we observed greater rejection in both conditions, compared to experiment 1. crucially, however, the difference remained between ‘the’ and ‘a’, with uses of ‘the’ being rejected more than uses of ‘a’ (65% vs. 41%, respectively). a mixed effect logistic regression revealed this difference to be significant (χ2(1)=4.7, p<.05). furthermore, as the right-hand side of figure 6 shows, participants almost always accepted the use of ‘the’ in the 1-adjective true trials, providing further evidence that the rejection of ‘the’ in experiments 1 and 2 was not driven by any potential infelicity associated with a presupposition failure. proceedings of elm 3: 144-151, 2025 andré eliatamby and lyn tieu: the role of definiteness in ad hoc implicatures. 149 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 6: mean acceptance of critical (underinformative) 1-adjective target trials and 1-adjective true trials for ‘a’ and ‘the’ conditions in experiment 2. 4. general discussion. the results from experiment 2 show that ad-hoc implicatures are more likely to be generated for sentences containing ‘the’ than for sentences containing ‘a’. ad-hoc implicatures for sentences containing the indefinite are attested; unlike in experiment 1, some participants always rejected the critical sentences containing ‘a’. nonetheless they are still generated less frequently with indefinite sentences compared to definite sentences. the particular determiner used in the sentence influences whether ad hoc implicatures are computed. if we assume that ‘the’ carries a uniqueness presupposition, a natural explanation for these results is that the requirement to satisfy uniqueness triggers implicature computation. such an explanation is difficult to support, however, if implicatures are only computed at the sentential level, as suggested by neo-gricean models of implicature (sauerland 2004, geurts 2010). the issue with generating ad hoc implicatures at the sentential level is that the denotation of “striped sweater” is the same in both the literal and enriched interpretation of sentences like (1), meaning that the uniqueness presupposition of ‘the’ is still not satisfied in the final enriched proposition. while a global ad hoc implicature yields the entailment that mary bought the sweater that was striped but not spotted, the set of “striped sweaters” still contains two sweaters. thus, under a neo-gricean framework of implicature generation, the need to satisfy the uniqueness presupposition of ‘the’ cannot be what is driving increased implicatures. uniqueness can be a trigger for implicature if implicatures are computed locally (chierchia et al. 2012, fox 2007) within the dp: (10) mary bought exh the striped sweater assuming that an alternative to ‘the striped sweater’ is ‘the striped and spotted sweater’, then the enriched dp generated in (10) is ‘the striped and not spotted sweater’. since there is only one sweater that is striped but not spotted, the uniqueness presupposition of (10) is therefore satisfied. the issue with this idea is that the exh operator attaches to propositional nodes, while definite proceedings of elm 3: 144-151, 2025 andré eliatamby and lyn tieu: the role of definiteness in ad hoc implicatures. 150 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ determiner phrases like ‘the striped sweater’ have been analyzed either as quantificational phrases or variables referring to some antecedently introduced referent; dps do not denote functions from propositions to truth conditions. furthermore, given their semantic type, the entailment relation involved in determining the set of alternatives used by exh would need to be redefined to capture the informativity relationship between dps. investigating the multiplicity inferences of plural definite noun phrases, mayr (2015) proposes an exhaustivity operator that applies at the predicate level, arguing that the definite article forces exhaustification to occur below it on the noun phrase directly. future research could explore how our findings concerning ad hoc implicatures might be accommodated within such a proposal, and within current theories of implicature more broadly. 5. conclusion. this study experimentally investigated the interaction between ad hoc implicatures and the definiteness property of english determiners. our linguistic intuition suggested that sentences with the english definite article ‘the’ trigger ad hoc implicatures more strongly than sentences with the indefinite ‘a’, even though standard models of implicature generation predict no difference. we confirmed this intuition in two truth value judgment tasks; participants were more likely to generate ad hoc implicatures with ‘the’ than ‘a’, although implicatures with ‘a’ were attested and the overall implicature generation rate was mediated by whether the stimuli were presented textually or auditorily. we have suggested that this result is hard to account for using neogricean models of implicature computation, and for localist models, may require changes to the distribution of exh or to the definition of the entailment relations involved in determining the sets of alternatives used by exh. references chierchia, g., fox, d., & spector, b. 2012. scalar implicature as a grammatical phenomenon. in handbücher zur sprach-und kommunikationswissenschaft/handbooks of linguistics and communication science semantics volume 3. de gruyter. https://library.oapen.org/handle/20.500.12657/23783. crain, s., & thornton, r. 2000. investigations in universal grammar: a guide to experiments on the acquisition of syntax and semantics. cambridge, ma: mit press. fox, d. 2007. free choice and the theory of scalar implicatures. in presupposition and implicature in compositional semantics pp. 71–120. springer. geurts, b. 2010. quantity implicatures. cambridge uk: cambridge university press. heim, i. (1991). artikel und definitheit [articles and definiteness]. in a. von stechow & d. wunderlich (eds.), semantik: ein internationales handbuch der zeitgenössischen forschung. hirschberg, j.l. 1991. a theory of scalar implicature. philadelphia pa: university of pennsylvania dissertation. mayr, clemens. 2015. plural definite nps presuppose multiplicity via embedded exhaustification. in d'antonio, sarah & mia wiegand (eds.). proceedings of semantics and linguistic theory (salt 25), 204-224. washington dc: linguistic society of america. https://doi.org/10.3765/salt.v25i0.3059 tomlinson, j. m., gotzner, n., & bott, l. 2017. intonation and pragmatic enrichment: how intonation constrains ad hoc scalar inferences. language and speech, 60(2), 200-223. https://doi.org/10.1177/0023830917716101. tomlinson, j. m., & ronderos, c. r. 2021. does intonation automatically strengthen scalar implicatures? semantics and pragmatics, 14(4), 1–30. https://doi.org/10.3765/sp.14.4. proceedings of elm 3: 144-151, 2025 andré eliatamby and lyn tieu: the role of definiteness in ad hoc implicatures. 151 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ disagreements do not automatically raise the standard of precision yifan wu & helena aparicio* abstract. speakers often choose to utter imprecise sentences that, albeit felicitous, are, strictly speaking, false (e.g., using ‘this bottle is empty’ to describe a bottle with a bit of water in it). the acceptability of an imprecise utterance hinges on the standard of precision (sop), a discourse parameter that governs how much imprecision is tolerated in a context. previous theoretical accounts (e.g., lewis 1979, klecha 2018) have argued that metalinguistic denials that target the assertability of an imprecise utterance (e.g., ‘no, this bottle is not empty!’) more or less force accommodation to a higher sop. the present study investigates the nature of this accommodation process. in particular, we ask whether metalinguistic disagreements result in an automatic update of the sop. in two acceptability judgment experiments, we show that imprecise utterances are not deemed unacceptable when embedded in a disagreement dialogue. our findings instead suggest that metalinguistic denials act as a request to raise the sop and that any potential updates ought to be signaled overtly in subsequent conversational moves. keywords. imprecision; standard of precision; metalinguistic disagreement; maximum standard absolute adjectives; discourse processing 1. introduction. during conversation, speakers often choose to stretch the boundaries of lexical representations by speaking loosely. for instance, in figure 1, alex describes the bottle as empty, despite being aware that it contains a small amount of water. such an instance showcases the phenomenon of imprecision, or loose talk (lewis 1979, lasersohn 1999, krifka 2002, 2007, kennedy 2007, syrett et al. 2010, lauer 2012, aparicio et al. 2015, leffel et al. 2016, aparicio terrasa 2017, klecha 2018, ronderos et al. 2024; a.o.), wherein a speaker chooses to utter a sentence that they judge to be pragmatically felicitous, despite it being strictly speaking false. alex: this bottle is empty. figure 1: imprecise description of a bottle containing some water. imprecision is pervasive in everyday communication, manifesting across a wide range of lexical items, such as maximum standard absolute adjectives (e.g., empty), verb phrases denoting events with incremental themes (e.g., peel the apple), or round numerals (e.g., one hundred), among others. predicates that can be subject to imprecision usually denote upper closed scalar meanings. here, we exemplify this core property of imprecision through the case of maximum standard absolute adjectives, which the current work uses as a testbed. maximum standard absolute adjectives have been argued to associate with adjectival scales (i.e., the scale encoding the dimension denoted by the adjective) that have an upper bound corresponding to the maximum degree on the relevant scale (e.g., the *we are grateful to the cornell lime lab, cornell interdisciplinary semantics group, cornell c. psyd, and the audiences at hsp2024 and elm3 for discussions. authors: yifan wu, cornell university (yw2578@cornell.edu) & helena aparicio, cornell university (haparicio@cornell.edu). proceedings of elm 3: 435-446, 2025 c©2025 yifan wu and helena aparicio published by the lsa with permission of the author(s) under a cc by license. 435 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ maximum degree of emptiness; rotstein & winter (2004), kennedy & mcnally (2005), kennedy (2007)), see figure 2.1 upper closed scale: figure 2: scale topology for maximum standard absolute adjectives. a precise interpretation of a maximum standard absolute adjective obtains when the adjective is predicated of an individual that instantiates the property to its maximum degree (e.g., a bottle that is completely empty). imprecise interpretations, on the other hand, obtain when the predicate is applied to an individual that bears the adjectival property to a non-maximal degree, as exemplified in figure 1. the felicity of an imprecise utterance has been shown to be contingent on features of the discourse context (van der henst et al. 2002, burnett 2014, aparicio terrasa 2017, ronderos et al. 2024). for instance, a movie theater with only three people sitting in it can be felicitously described as empty in the context of a highly anticipated movie premiere that was expected to be a blockbuster. in this case, the imprecise use of the predicate is licensed: three people is a small enough number for it to be safely ignored. however, in a higher-stakes context such as an emergency fire evacuation, the same imprecise interpretation of the predicate stops being available, since ignoring the three people sitting in the theater could have fatal consequences.2 additionally, recent studies have shown that comprehenders take into account information about the speaker’s goals (mathis & papafragou 2022), as well as social information about the speaker’s identity (beltrama & schwarz 2021, 2022, 2024) to determine whether to adopt an imprecise interpretation of the utterance. here, we take the acceptability of an imprecise utterance to be determined by the standard of precision (sop, lewis 1979), a latent discourse parameter—or in lewis’ terms, an element of the conversational score—that governs the degree of imprecision tolerated in a given discourse. because the sop is not observable, conversational agents must rely on contextual cues (some of which have already been mentioned above) to infer likely parametrizations of the sop. while in most cases speakers and listeners successfully align on the value of the sop, this implicit coordination process is not infallible; occasionally, interlocutors assume conflicting parametrizations that can eventually cause conversational disruptions. of particular interest to us are instances where the speaker asserts an utterance whose felicity is contingent on a sufficiently low sop, whereas the listener has all along been assuming a stricter one. in such situations, the listener may choose to go along with the speaker’s conversational move by updating their beliefs about the sop to a sufficiently low value so as to render the speaker’s utterance felicitous. alternatively, the listener might be unwilling to accommodate. in such cases, the only discourse move available to the listener is to object to the assertability of the utterance through a metalinguistic denial or disagreement (horn 1989, barker 2002, 2013), as shown in (1). the disagreement in (1) is therefore not about the factual state of the world, both alex and andy acknowledge that the bottle contains some water. 1while maximum adjectives minimally have upper bounds, adjectives like empty actually associate with fully closed scales, i.e., scales that are closed on both the upper and the lower end. antonymous pairs of gradable adjectives map their arguments onto the same scale but impose inverse orderings on their shared domains. for the antonym pair full/empty, both maximum standard absolute adjectives, the maximum of full is empty’s minimum and vice versa. 2this example is based on an example provided by burnett (2014). proceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 436 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ rather, andy is taking issue with the appropriateness of the predicate empty as a suitable descriptor of the referent. put differently, andy is indirectly challenging the low sop assumed by alex and signaling that a higher sop should be adopted.3 (1) a. alex: this bottle is empty. [uttered as a description of the bottle in figure 1] b. andy: no, this bottle is not empty; there’s a bit of water in it! previous authors have observed that metalinguistic denials such as (1-b) are hard to resist (klecha 2018: 92), more or less forcing the listener, in our running example alex, to accommodate to a higher sop (lewis 1979, lauer 2012, klecha 2018). here we investigate whether this need to accommodate is a by-product of the discursive import of metalinguistic denials. more specifically, we ask whether metalinguistic denials such as (1-b) lead to an automatic update of the sop. we consider two hypotheses. hypothesis 1 states that challenging the sop through a metalinguistic denial automatically updates this discourse parameter, thereby superseding previous parametrizations. hypothesis 1 makes the prediction that disagreements should decrease the acceptability of a previous imprecise utterance, since after the challenge only the higher sop should be operative. contra hypothesis 1, hypothesis 2 posits that metalinguistic denials act as a request to raise the sop, but do not directly update it. this hypothesis therefore predicts that the looser sop can remain operative after the challenge and that the acceptability of the original imprecise utterance should not decrease. we report results from two acceptability judgment studies, where we find evidence for hypothesis 2. our results indicate that imprecise utterances are not deemed unacceptable in disagreement dialogues, suggesting that the lower sop remains operative even after being challenged. this implies that any potential updates to the sop ought to occur in subsequent conversational moves. the remainder of this paper proceeds as follows. section 2 presents the results pertaining to experiment 1. in section 3, we present experiment 2 and the comparison analysis between the two experiments. finally, section 4 provides a general discussion of our findings and section 5 concludes the paper. 2. experiment 1. the goal of experiment 1 was to obtain interpretational preferences for imprecise utterances in isolation. results pertaining to experiment 1 were later used as a baseline for comparison with results from experiment 2, in which the same visual stimuli were paired with disagreement dialogues (see section 3). 2.1. materials & design. we constructed 24 five-point scales instantiating different maximum standard absolute properties to varying degrees (see figure 3). individual scale-points were combined with a written statement of the form ‘this [object] is [adjective]’ (e.g., this bottle is empty), where the noun was always an appropriate descriptor of the depicted object and the adjective matched the property represented in the scale the picture was part of (see left panel of figure 5). the five scale-points pertaining to each of the 24 scales were distributed across five lists following a latin-square design. this ensured that each participant judged one single scale-point 3metalinguistic disagreements over imprecise utterances have been argued to be unidirectional (e.g., lewis 1979, klecha 2018), meaning that such implicit challenges can be used to raise the sop, but not to lower it. in klecha’s terms, the sop can be raised incidentally to the content of an expression. however, in order to lower the sop, klecha argues that speakers must engage in an explicit metalinguistic negotiation (klecha 2018: 93). proceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 437 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ per item. trial order was randomized within each list and for each participant. figure 3: five scale-points corresponding to the scale ‘empty bottle’. 2.1.1. norming study. the 24 scales tested in experiment 1 were normed in order to ensure that the lower scale points (s1-s4) tolerated some non-negligible amount of imprecision. in each trial, the full five-point scale was paired with a statement of the form ‘this [object] is [adjective]’. as in experiment 1, the noun always matched the depicted object and the adjective encoded the property represented in the visual scale. the study consisted of a forced-choice picture-matching task in which participants were instructed to assess whether the statement was an appropriate description of each individual scale-point. unlike in experiment 1, the full scale was available to participants at the moment of providing their judgements. participants gave their answers by choosing one of three possible responses (i.e., ‘yes’, ‘no’ or ‘unsure’) for each of the five points in the scale (see left panel of figure 4). thirty adult native speakers of american english recruited through the crowd-sourcing platform prolific participated in the study. participants were compensated at a rate of $15 per hour. results are shown in the right panel of figure 4. as can be observed in the plot, precise interpretations were preferred over imprecise ones. this is shown by the fact that the precise scale-point (s5) received the highest amount of ‘yes’ responses (green bars). more important for us, participants demonstrated some degree of tolerance for imprecise interpretations in all the lower scale-points, as indexed by the ‘yes’ responses in s1-4. the acceptability of imprecise interpretations, however, was gradient: the further the scale point strayed from the endpoint-oriented interpretation, the less available imprecise interpretations became. figure 4: left: norming study item example; right: norming study results. 2.2. procedure. the experiment was administered remotely through the pcibex farm platform (zehr & schwarz 2018). at the beginning of the experiment, participants provided informed consent, completed a demographic questionnaire, and engaged in three practice trials designed to acclimate them to the experimental setup and response protocol. in the main part of the experiproceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 438 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ment, participants saw a visual stimulus accompanied by a sentence (e.g., this bottle is empty, see left panel of figure 5). the experimental task consisted of a forced-choice picture-matching task, where participants were asked to judge whether the written statement was an appropriate description of the picture. participants provided their answer by choosing one of three possible options: ‘yes’, ‘no’, and ‘unsure’. 2.3. participants. thirty participants were recruited through the web platform prolific and were compensated at a rate of $15/hour. all participants were native speakers of american english and were at least 18 years old. figure 5: left: experiment 1 item example; right: experiment 1 results. 2.4. results & discussion. results are shown in figure 5. experiment 1 qualitatively replicates the results obtained in the norming study. first, participants continue to display a strong preference for precision over imprecision. to confirm this statistically, we fit a logistic mixedeffects regression model to the binarized response variable (i.e., yes-responses were coded as 1, and noand unsure-responses were coded as 0), using scale point as a fixed effect with s5 as the reference level. the model also included random intercepts by-item and by-participant, as well as by-condition random slopes. model outputs, shown in table 1a, reveal significant effects for all comparisons (all p’s < 0.001), confirming that the precise scale-point (s5) received a significantly higher proportion of yes-responses when compared to those scale-points that only supported an imprecise interpretation. second, all lower scale-points (s1-s4) exhibit some tolerance for imprecision (see the green bars corresponding to yes-responses in the right panel of figure 5). as in the norming study, tolerance for imprecision follows a gradient: the closer the scale-point is to the maximum scalar degree, the more available imprecise interpretations become. among the scale-points that were incompatible with a precise interpretation (i.e., s1-s4), s1 showed the lowest proportion of yesresponses (20%), whereas s4 showed the highest (30%). to further examine this gradient effect, we recoded our categorical predictor scale point using a forward-difference contrast scheme (a coding system that compares the mean of the dependent variable for one level of a categorical variable to the mean of the next level). none of the comparisons reached significance (see table 1b). finally, we note that participants displayed very little uncertainty regarding the appropriateness of the predicates, as reflected by the low proportion of unsure-responses (all scale-points displayed proceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 439 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ proportions lower than 5%). table 1: model outputs for experiment 1. (a) model 1: preference for precision analysis. β se z p (intercept) 2.81 0.44 6.44 1.18e-10*** s5 vs. s1 -5.85 0.86 -6.84 7.81e-12*** s5 vs. s2 -6.00 0.90 -6.67 2.56e-11*** s5 vs. s3 -5.53 0.88 -6.28 3.47e-10*** s5 vs. s4 -5.11 0.92 -5.56 2.70e-08*** (b) model 2: gradience analysis. β se z p (intercept) -2.99 0.64 -4.69 2.68e-06*** s1-s2 -0.19 0.36 -0.53 0.5934 s2-s3 0.27 0.35 0.78 0.4368 s3-s4 0.71 0.38 1.86 0.0629. significance levels: <.001*** <.01** <.05* <.1 taken together, experiment 1 results show that judgments elicited by presenting the scalepoints in isolation are qualitatively comparable to the norming study results, where the same items were judged as part of the five-point scale: overall, precision was preferred over imprecision, but all the lower scale-points allowed for some degree of imprecision, especially scale-points that were closer to the endpoint. more importantly, results from experiment 1 provide us with a baseline for comparison with results pertaining to experiment 2, to be presented in the next section. 3. experiment 2. the goal of experiment 2 was to determine whether the acceptability of an imprecise utterance declines when judged as part of a disagreement dialogue. 3.1. materials & design. experiment 2 followed the same design and tested the same stimuli as experiment 1, with one crucial modification: the visual stimuli were paired with a disagreement dialogue. all the disagreements consisted of a speaker assertion of the form ‘this [object] is [adjective]’ (e.g., ‘this bottle is empty’) followed by a metalinguistic denial of the form ‘no, this [object] is not [adjective]’ (e.g., ‘no, this bottle is not empty’, see left panel of figure 7). additionally, twenty-four fillers were included. in the filler trials, participants were presented with a visual stimulus coupled with a disagreement dialogue. however, unlike experimental trials, the predicates under discussion were not maximum standard adjectives but rather color adjectives (e.g., yellow) or minimum standard gradable adjectives (e.g., dashed). in half of the filler trials, the image clearly matched both the adjectival property and the noun included in the first speaker’s utterance, whereas in the other half of the filler trials the visual stimulus could not be described with the adjectival property under discussion, only with the noun (see figure 6). experimental materials corresponding to the critical trials were distributed in five lists following a latin square design. each list was complemented with the 24 filler trials. each participant was randomly assigned to a list and the order of the 48 trials within each list were randomized for each participant. alex: this fish is yellow. andy: no, this fish is not yellow. alex: this circle is dashed. andy: no, this circle is not dashed. figure 6: left: color filler example; right: minimum standard absolute adjective filler example. proceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 440 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3.2. procedure. experiment 2 followed the same procedure used in experiment 1 with one important distinction. in experiment 2, participants’ task was to judge whether only one of the speakers was right, or whether both of them could be right by selecting one of three options: ‘only the first speaker is right’, ‘only the second speaker is right’, or ‘both of them can be right’ (henceforth, ‘first’, ‘second’ and ‘both’; see left panel of figure 7). in scale-points s1-4, which are only compatible with an imprecise interpretation of the first speaker’s utterance, a first-response indicates that the imprecise utterance was judged to be acceptable, even after having been targeted by a metalinguistic denial. first-responses therefore entail that the lower sop remains operative, even after the second speaker takes issue with the assertability of the imprecise utterance. in contrast, second-responses indicate that the imprecise utterance was deemed unacceptable. this response therefore indexes alignment with the higher sop adopted by the second speaker. finally, a bothresponse indicates that the disagreement was judged to be faultless (kölbel 2004, barker 2013, kennedy 2013, kaiser & rudin 2020, 2021, pecsok & aparicio 2024). faultless disagreements are a type of disagreement where neither party in the discourse is judged to be at fault. we take this type of answer to be compatible with a lower sop. 3.3. participants. participants consisted of 60 native speakers of american english who were at least 18 years old. all participants were recruited through the web platform prolific. participation was compensated at a rate of $15 per hour. two participants were removed from data analysis due to failure to reach a 90% accuracy threshold in the filler trials. 3.4. predictions. as discussed in section 1, our first hypothesis (h1) states that metalinguistic denials automatically raise the sop. h1 therefore predicts that the acceptability of the imprecise utterance should decrease in experiment 2 compared to experiment 1, where the same utterance is judged in isolation. more specifically, in experiment 2, s1-4 should display a substantial increase in second-responses compared to the proportion of no-responses observed in experiment 1. crucially the increase in second-responses should be accompanied by a decrease in first-responses, such that the proportion of first-responses in experiment 2 should be lower compared to the proportion of yes-responses in experiment 1. no differences are predicted for s5. our second hypothesis (h2) posits that metalinguistic challenges only act as a request to update the sop. h2 therefore predicts that the lower sop should remain operative after a metalinguistic denial, and that the acceptability of the imprecise utterance should not be negatively impacted by the disagreement. under this hypothesis, we expect comparable acceptability rates for imprecise utterances across experiments 1 and 2. this pattern should materialize as comparable rates of yes/first-responses and no/second-responses respectively across experiments 1 and 2. alternatively, if participants select both-responses at a high rate in experiment 2, we would expect this choice to be at the expense of second-responses. this should result in lower selection rates of second-responses in experiment 2, compared to no-responses in experiment 1. 3.5. results & discussion. results for experiment 2 are presented in the right panel of figure 7. we note that the preference for precision remains evident in experiment 2, as shown by the high proportion of first-responses (green bars) in s5 compared to all other scale-points. to statistically confirm this preference, we fit a mixed effects logistic regression model to the binarized response data (i.e., first-responses were coded as 1 and bothand second-responses as 0). this dependent proceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 441 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ variable was predicted from scale point, with s5 as the reference level. the model also included by-item and by-participant random intercepts, as well as random slopes by scale-point for both participants and items. model outputs revealed significant differences in all comparisons (see table 2a). alex: this bottle is empty. andy: no, this bottle is not empty. figure 7: left: experiment 2 item example; right: experiment 2 results. results also reveal that participants selected both at a higher rate in s1-s4 compared to s5. to statistically assess this qualitative pattern, we binarized our response data, such that both-responses were coded as 1, while the remaining two options (first and second) were coded as 0. a logistic mixed effects regression model was fit to this new binomial variable using scale point as a fixed effect. s5 was coded as the reference level. random intercepts by items and participants were also included. no random slopes were added to the model due to convergence issues. as shown in table 2b, all comparisons reached significance. table 2: model outputs for experiment 2. (a) model 1: preference for precision analysis. β se z p (intercept) 1.48 0.19 7.83 5.00e-15*** s5 vs. s1 -5.56 1.03 -5.41 6.24e-08*** s5 vs. s2 -5.99 1.32 -4.55 5.28e-06*** s5 vs. s3 -5.22 1.00 -5.22 1.77e-07*** s5 vs. s4 -3.50 0.61 -5.79 7.18e-09*** (b) model 2: both-responses analysis. β se z p (intercept) -3.10 0.38 -8.21 2.26e-16*** s5 vs. s1 1.35 0.28 4.89 1.01e-06*** s5 vs. s2 1.19 0.28 4.28 1.90e-05*** s5 vs. s3 1.58 0.28 5.74 9.24e-09*** s5 vs. s4 1.36 0.28 4.93 8.27e-07*** in order to determine whether the predictions of our two hypotheses (see section 3.4) are borne out, we conducted comparison analyses between experiment 1 and experiment 2. specifically, we compared yes-responses in experiment 1 to first-responses in experiment 2, as these responses indicate that participants took the lower standard to be operative. additionally, we compared noresponses in experiment 1 to second-responses in experiment 2, as these responses indicate that participants rejected the imprecise utterances in s1-s4, and therefore did not take the lower standard to be operative. we did not perform a comparison between the unsureand both-responses in experiments 1 and 2 respectively, as these responses are not directly comparable. four new binary variables were constructed. for experiment 2, first-responses were coded as 1 and secondand both-responses were coded as 0. this variable reflected the acceptability of the imprecise utterance in s1-s4. a second variable in which second-responses were coded as 1 and proceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 442 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ all other options as 0 was also created. this second random variable reflected the unacceptability of the imprecise utterance in s1-s4. the same procedure was followed for experiment 1 with yesversus no-responses, with unsure-responses always being coded as 0. the four binomial variables were appended and coded based on 1) whether the observation belonged to experiment 1 or 2. we refer to this factor as experiment; and 2) whether the imprecise utterance was accepted (i.e., yes in experiment 1, and first in experiment 2), or rejected (i.e., no in experiment 1, and second in experiment 2). we refer to this factor as acceptability. response proportions by item, scalepoint and experiment were obtained within each acceptability level. this summarized variable was used as the dependent measure in all subsequent analyses. a series of linear mixed effects models were fit to the data pertaining to each scale-point, with experiment, acceptability and their interaction as fixed effects, and random intercepts and slopes by item scale. at all scalepoints, we found a significant interaction effect between experiment and acceptability (s1: β = −0.12, t = −3.04, p < 0.01; s2: β = −0.15, t = −5.07, p < 0.001; s3: β = −0.13, t = −4.25, p < 0.001; s4: β = −0.13, t = −3.64, p < 0.001; s5: β = 0.07, t = 2.87, p < 0.01).4 *** *** *** *** n.s. (a) imprecise utterance rejected. n.s. n.s. * n.s. * (b) imprecise utterance accepted. figure 8: experiments 1 & 2 comparison. in order to further probe the interactions, we conducted simple effects analyses for each scalepoint with experiment as the independent factor and by-item random intercept and slopes. this analysis is visualized in figure 8. results show that proportions of responses rejecting the lower standard (left panel in figure 8) were significantly lower in experiment 2 compared to experiment 1 in all the lower scale-points (s1: β = −0.16, t = −4.02, p < 0.001; s2: β = −0.17, t = −5.99, p < 0.001; s3: β = −0.19, t = −5.14, p < 0.001; s4: β = −0.17, t = −4.05, p < 0.001). no significant difference was observed at the precise scale point (s5: β = 0.01, t = 0.65, p > 0.1). we now move on to the analysis of responses indicating participants’ acceptance of the lower standard (right panel of figure 8b). no significant differences were found between yesand first-responses in most of the lower scale points (s1: β = −0.04, t = −1.35, p > 0.1; s2: β = −0.02, t = −0.67, p > 0.1; s4: β = −0.04, t = −1.74, p > 0.05), although s3 and s5 did reach significance (s3: β = −0.05, t = −2.74, p < 0.05; s5: β = −0.06, t = −2.69, p < 0.05). comparison analyses of experiments 1 and 2 show that proportions of first-responses in ex4due to space constraints, we do not include model estimates and significance levels for the main effects. however, we note that such comparisons do not have any bearing on the hypotheses under consideration. proceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 443 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ periment 2 and yes-responses in experiment 1 were comparable across scale-points s1-s4, with the exception of s3, where the proportion of first-responses was significantly lower. conversely, the proportions of second-responses in experiment 2—a choice that is only compatible with a higher sop—were systematically lower than no-responses in experiment 1 across all the imprecise scalepoints (s1-s4). this decrease in selection rates of second-responses was therefore incurred by the high selection rates observed for both-responses, a point that we return to in the general discussion. overall, our findings suggest that the acceptability of the imprecise utterance remains largely stable across the two experiments, a finding that can be better accommodated by h2, which states that metalinguistic denials act only as a request to raise the sop. importantly, we also find that rejections of the imprecise utterance decrease across the two experiments. this second finding argues against h1, which predicts higher rejection rates of the imprecise utterance after a metalinguistic denial. taken together, our results therefore allow us to conclude that challenging the assertability of an imprecise utterance does not incur an automatic update of the sop. 4. general discussion. in this section, we discuss our findings in the larger context of the literature on imprecision. we first note that both experiments 1 & 2, as well as our norming study, showed a clear preference for precise interpretations over imprecise ones. these results replicate previous findings showing that precise interpretations are not only judged to be more acceptable, but are also processed faster (syrett et al. 2010, aparicio et al. 2015, aparicio terrasa 2017, leffel et al. 2017, ronderos et al. 2024). we now turn to the research question addressed in the current paper. our starting point was the observation that the move to raise the sop by means of a metalinguistic denial is hard to resist (lewis 1979, klecha 2018). our first hypothesis (h1) proposed that this difficulty is a direct consequence of the discursive import of metalinguistic denials. in particular, h1 states that denials effectively raise the sop. as has been discussed, our results argue against this view: the acceptability of an imprecise utterance overall was not negatively affected when embedded in a disagreement dialogue. our results can be better accounted for by the second hypothesis under consideration (h2), which states that metalinguistic denials act only as a request to raise the sop. it is important to note that our findings do not argue against previous claims that metalinguistic denials more or less cause the listener to accommodate to a higher sop. our results however suggest that, to the extent that this accommodation takes place, it ought to be overtly signaled in a subsequent conversational move, such as a concession or a retraction. in this respect, accommodating to a higher sop differs from better understood types of accommodation, such as presupposition accommodation (cf. klecha (2018) for a similar observation). if h2 is on the right track, and metalinguistic denials do indeed act as a request to raise the sop, it remains an open question what the specific effect of this move is on the sop. one possibility is that the lower sop remains operative until the listener either overtly agrees to adopt a raised threshold, or further challenges it. a second possibility is that the denial has the effect of suspending the sop, such that it is temporarily undefined until the subsequent conversational move. this view is partially supported by one aspect of our findings, namely the high rates of both-responses obtained in s1-4 in experiment 2. the fact that in this experiment participants judged the disagreement to be faultless at high rates suggests that they took both the lower and the higher standard, to be operative. one possible way of reconciling these two incompatible proceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 444 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ parametrizations would be to assume that as long as the two speakers have incompatible beliefs about the sop, this parameter ought to be undefined in the common ground. 5. conclusion. this paper investigates the discourse dynamics of imprecision, focusing on whether metalinguistic disagreements affect the acceptability of imprecise utterances. building on previous claims that such disagreements trigger accommodation to a higher standard of precision (sop) (lewis 1979, klecha 2018), we tested two hypotheses. hypothesis 1 (h1), consistent with prior work, proposes that metalinguistic denials automatically raise the sop, rendering the original imprecise utterance unacceptable. hypothesis 2 (h2) suggests that metalinguistic denials serve as a request to raise the sop rather than enforcing an update, allowing the lower sop to remain operative. our results show that imprecise utterances remain as acceptable after a metalinguistic denial as they are in isolation, in line with h2. this indicates that lower sops persist even after metalinguistic challenges, with disagreements prompting but not mandating higher sop adoption. in ongoing work, we investigate whether and how the discourse commitments (lauer 2012) incurred by subsequent conversational moves (e.g., concessions vs. retractions) update the sop. references aparicio, helena, ming xiang & christopher kennedy. 2015. processing gradable adjectives in context: a visual world study. semantics and linguistic theory (salt) 25. 413–432. https://doi.org/10.3765/salt.v25i0.3128. aparicio terrasa, helena. 2017. processing context-sensitive expressions: the case of gradable adjectives and numerals: the university of chicago dissertation. barker, chris. 2002. the dynamics of vagueness. linguistics and philosophy 25(1). 1–36. barker, chris. 2013. negotiating taste. inquiry 56(2-3). 240–257. beltrama, andrea & florian schwarz. 2021. imprecision, personae, and pragmatic reasoning. semantics and linguistic theory (salt) 31. 122–144. https://doi.org/10.3765/salt.v31i0.5107. beltrama, andrea & florian schwarz. 2022. social identity, precision and charity: when less precise speakers are held to stricter standard. semantics and linguistic theory (salt) 32. 575–598. https://doi.org/10.3765/salt.v1i0.5406. beltrama, andrea & florian schwarz. 2024. social identity affects imprecision resolution across different tasks. semantics and pragmatics 17. 10–ea. https://doi.org/10.3765/sp.17.10. burnett, heather. 2014. a delineation solution to the puzzles of absolute adjectives. linguistics and philosophy 37. 1–39. https://doi.org/10.1007/s10988-014-9145-9. horn, laurence r. 1989. a natural history of negation. chicago: university of chicago press. kaiser, elsi & deniz rudin. 2020. when faultless disagreement is not so faultless: what widelyheld opinions can tell us about subjective adjectives. proceedings of the linguistic society of america 5(1). 698–707. https://doi.org/10.3765/plsa.v5i1.4757. kaiser, elsi & deniz rudin. 2021. arguing with experts: subjective disagreements on matters of taste. proceedings of the annual meeting of the cognitive science society 43(43). 924–930. kennedy, christopher. 2007. vagueness and grammar: the semantics of relative and absolute gradable adjectives. linguistics and philosophy 30. 1–45. https://doi.org/10.1007/s10988006-9008-0. proceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 445 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ kennedy, christopher. 2013. two sources of subjectivity: qualitative assessment and dimensional uncertainty. inquiry 56(2–3). 258––277. https://doi.org/10.1080/0020174x.2013.784483. kennedy, christopher & louise mcnally. 2005. scale structure, degree modification, and the semantics of gradable predicates. language 81(2). 345–381. klecha, peter. 2018. on unidirectionality in precisification. linguistics and philosophy 41. 87– 124. https://doi.org/10.1007/s10988-017-9216-9. krifka, manfred. 2002. be brief and vague! and how bidirectional optimality theory allows for verbosity and precision. in david restle & dietmar zaefferer (eds.), sounds and systems: studies in structure and change. a festschrift for theo vennemann, 439–458. berlin, new york: de gruyter mouton. https://doi.org/10.1515/9783110894653.439. krifka, manfred. 2007. approximate interpretation of number words. berlin: humboldtuniversität zu berlin, philosophische fakultät ii. https://doi.org/10.18452/9508. kölbel, max. 2004. faultless disagreement. proceedings of the aristotelian society 104(1). 53–73. https://doi.org/10.1111/j.0066-7373.2004.00081.x. lasersohn, peter. 1999. pragmatic halos. language 75(3). 522–551. lauer, sven. 2012. on the pragmatics of pragmatic slack. proceedings of sinn und bedeutung 16(2). 389–402. leffel, timothy, ming xiang & christopher kennedy. 2016. imprecision is pragmatic: evidence from referential processing. semantics and linguistic theory (salt) 26. 836–854. https://doi.org/10.3765/salt.v26i0.3937. leffel, timothy, ming xiang & christopher kennedy. 2017. interpreting gradable adjectives in context: domain distribution vs. scalar representation. unpublished manuscript. lewis, david. 1979. scorekeeping in a language game. journal of philosophical logic 8. 339–359. https://doi.org/10.1007/bf00258436. mathis, ariel & anna papafragou. 2022. agents’ goals affect construal of event endpoints. journal of memory and language 127. 104373. https://doi.org/10.1016/j.jml.2022.104373. pecsok, emily & helena aparicio. 2024. how can they both be right?: faultless disagreement and semantic adaptation. proceedings of the annual meeting of the cognitive science society 46. 3939–3945. ronderos, camilo r., ira noveck & ingrid lossium falkum. 2024. straight enough: deriving imprecise interpretations of maximum standard absolute adjectives. glossa psycholinguistics 3(1). 1–36. https://doi.org/10.5070/g60111411. rotstein, carmen & yoad winter. 2004. total adjectives vs. partial adjectives: scale structure and higher-order modifiers. natural language semantics 12. 259–288. syrett, kristen, christopher kennedy & jeffrey lidz. 2010. meaning and context in children’s understanding of gradable adjectives. journal of semantics 27(1). 1–35. https://doi.org/10.1093/jos/ffp011. van der henst, jean–baptiste, laure carles & dan sperber. 2002. truthfulness and relevance in telling the time. mind & language 17(5). 457–466. https://doi.org/10.1111/1468-0017.00207. zehr, jeremy & florian schwarz. 2018. penncontroller for internet based experiments (ibex). https://doi.org/10.17605/osf.io/md832. proceedings of elm 3: 435-446, 2025 yifan wu and helena aparicio: disagreements do not automatically raise the standard of precision. 446 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ mechanistic support language in colombian spanish-speakers jennifer barbosa1, paola pinzón-henao1, angelina pasquella, paul muentener & laura lakusta* abstract. beyond basic spatial relations (e.g., teddy on table), we know little about how children learn to talk about mechanical support events (e.g., objects attached/hung from a surface via tape) and map them onto linguistic structures. moreso, the majority of the research that has been done focuses on children learning english—a language that has several verbs that lexicalize support via a specific mechanism (levin 1993; e.g., glue, tape, clip, etc.). the current study seeks to deepen our understanding of spatial language acquisition by diversifying the populations that have been studied. specifically, 4-to-6 year-old monolingual spanish-speaking children and adults in colombia viewed mechanical support events (e.g., girl puts paper on door via tape) and were then asked, ‘can you tell me what my sister did with my toy?’. both children and adults used non-mechanism (e.g., poner = ‘put’, colgar = ‘hang’) and mechanism verbs (e.g., pegar = ‘stick’); the use of mechanism verbs increased from 4-to-6 years of age. in addition, whether the mechanism was visible in the event influenced how it was mapped to language; when the mechanism was visible (vs. when it was hidden), children and adults were more likely to encode the mechanism in a prepositional phrase (e.g., lo colgó con un gancho = ‘she hung it with a clip’). these findings shed light on the development of mechanical support language in spanish-speaking children, the influence of context—specifically, visibility of mechanism—on language, as well as the lexicalization patterns for encoding physical support in spanish more generally. keywords. support relations; mechanical support; cognition; spanish language; language development; force dynamics 1. introduction. children’s acquisition of spatial language begins at an early age (johnston & slobin 1979) and has been shown to play a role in children’s later academic success, especially stem-related disciplines, such as science and math (zimmerman et al. 2018). yet we still know relatively little about how children acquire spatial language, and how acquisition may differ based on the language being learned. to explore this, we focus on the spatial domain of physical support. physical support can be understood as the causal-force dynamic relations between objects in which one object prevents another object from falling (coventry et al. 1994, herskovits 1986, landau 2020, vandeloise 1991). the types of causal-force dynamic relations comprising physical support events are quite broad: all support relations involve some knowledge of gravity, some support relations require knowledge that solid objects can’t pass through each other (e.g., teddy on top of * we would like to acknowledge the three private schools and their personnel in manizales, colombia (colegio lions bellavista, jardín infantil mentes maravillosas, and colegio seminario menor de nuestra señora del rosario) that allowed data collection to occur during school hours. we would also like to acknowledge valentina cano, an undergraduate research assistant and native spanish speaker, who assisted with spanish transcriptions of response utterances, and cristina manteca gaucho, also a native spanish speaker and doctoral student with expertise in hispanic linguistics, who helped with interlinear gloss examples in this manuscript. authors: jennifer barbosa, montclair state university (barbosaj4@montlcair.edu), paola pinzón-henao, montclair state university (pinzonhenaop1@montclair.edu), angelina pasquella, montclair state university (pasquellaa1@montclair.edu), paul muentener, tufts university (paul.muentener@tufts.edu), & laura lakusta, montclair state university (lakustal@montclair.edu). jennifer barbosa and paola pinzón-henao contributed equally as first authors. proceedings of elm 3: 32-42, 2025 c©2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta published by the lsa with permission of the author(s) under a cc by license. 32 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ table), while others require mechanism-specific knowledge (e.g., hooks, tape, and velcro all have different properties that prevent objects from falling). in addition, the cause of the support can be visible or hidden (e.g., picture taped to a wall where the viewer can see the tape, or it is hidden behind the figure object; see figure 1). figure 1: examples of stimuli 1.1. encoding physical support in language. the way in which language maps to physical support relations is complex—with language differentiating the semantic space of support into, at least, two distinct types—support-from-below (sfb) and mechanical support. levinson and wilkins (2006) report that for many languages, there is a basic locative construction (blc) that maps canonically to sfb. in spanish, the blc is estar en (‘to be on’) which encodes static support configurations as in (1). the blc describes the state of a support event (i.e., once the object is already attached to the other) and does not explain how the causal change of state between the two objects occurred (i.e., from not having contact to being attached to the other). (1) el oso est-á en la mesa def.m bear be.3sg.pst on def.f table ‘the bear is on the table.’ for dynamic events of support, the verb poner (‘put’) acts semantically similarly to the blc for static events—it is semantically empty in that it does not encode details about the support, such as the mechanism that was used (picture on the wall via a nail) or the result orientation of the figure object (picture hanging on a wall, which means it is oriented downward). poner differs from estar in that it encodes the action which results in the static configuration as in (2). (2) ella pus-o la foto en la pared she put. 3sg.pst def.f picture on def.f wall ‘she put the picture on the wall.’ proceedings of elm 3: 32-42, 2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta: mechanistic support language in colombian spanish-speakers. 33 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ in contrast to the blc that maps to sfb (levinson & wilkins 2006), there is no single linguistic construction that has been proposed to map canonically to mechanical support. rather, there are a variety of linguistic constructions. 1.2. the language of mechanical support. mechanical support may be encoded with the blc or a simple verb as in (3) below: (3) a. estar en be.inf on ‘be on’ b. la foto est-á en la pared def.f picture be.3sg.pst on def.f wall ‘the picture is on the wall.’ c. ella la pus-o en la pared she it.f put. .3sg.pst on def.f wall ‘she put it on the wall.’ mechanical support can also be encoded by a variety of different lexical verbs in spanish. importantly, lexical verbs vary as to what component of the mechanical support is encoded. in (4a) colgar (‘hang’) encodes the orientation, but not the specific mechanism of attachment (i.e., the paper is oriented downward from the tree); in (4b) pegar (‘stick’) encodes a characteristic/property of the mechanism (i.e., sticky); and in (4c) the specific mechanism of support (fastener) is encoded (i.e., clavar = ‘nail’). of note, in english, lexical verbs that encode the mechanism (often denominals) are common in mechanical support language (e.g., tape, pin, clip) yet less is known about whether they are used in spanish. (4) a. la niña colg-ó́ el papel del árbol def.f girl hang.3sg.pst def.m paper of+def.m tree ‘the girl hung the paper on the tree.’ b. la niña peg-ó el papel al árbol def.f girl stick.3sg.pst def.m paper to+def.m tree ‘the girl stuck the paper to the tree.’ c. la niña clav-ó el papel al árbol def.f girl nail. .3sg.pst def.m paper to+def.m tree ‘the girl nailed the paper to the tree.’ since verbs such as (3c) poner and (4a) colgar do not encode the mechanism, we refer to these types of verbs as non-mechanism verbs. in contrast, since verbs such as (4b) pegar and (4c) clavar encode the mechanism (a property or the actual fastener), we refer to these types of verbs as mechanism verbs. this classification was motivated by levin’s (1993) english classification of verbs, which characterizes verb classes according to their semantics. linguistic structures other than verbs can be used to encode the mechanism, such as prepositional phrases, (5a) (con una cinta = ‘with a tape’). speakers have a variety of options in how they combine phrases. for example, speakers may combine a simple verb (5a) (poner en = ‘put on’) or another lexical verb (5b, 5c) with a prepositional phrase. these events can also be described by combining two separate clauses (5d). proceedings of elm 3: 32-42, 2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta: mechanistic support language in colombian spanish-speakers. 34 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (5) a. ella lo pus-o en el tablero con una cinta she it.m put.3sg.pst on def.m board with indf.f tape ‘she put it on the board with a piece of tape.’ b. ella lo colg-ó con un gancho she it.m hang.3sgpst with indf.m clip ‘she hung it with a clip.’ c. ella lo colg-ó de un gancho she it.m hang.3sgpst from indf.m clip ‘she hung it from a clip.’ d. ella le pus-o cinta y lo pus-o en la puerta she 3sg.io put.3sg.pst tape and it.m put.3sg.pst on def.f door ‘she put tape on it and put it on the door.’ we know from studies of other domains (e.g., manner of motion, causation) that languages vary in their lexicalization patterns (see talmy 2000). given this, in addition to the range of linguistic constructions that are offered to map to mechanical support, in the current study we ask 1) how do monolingual spanish speakers encode dynamic mechanical support events? and 2) how may these descriptions change over development in monolingual spanish speakers? in addition, since the mechanism of support can be visible to the viewer or hidden in mechanical support events (see figure 1), we vary this feature in the current study. research suggests that children reason about the visible and hidden properties that explain how objects behave and interact (schulz 2012); thus, in the current study we ask how the visibility of the mechanism may affect how language maps to mechanical support events. 2. participants. since preschool-aged spanish-speaking children in the u.s. are exposed to english in various contexts outside of the home (e.g., preschool settings, extracurricular activities; welsh & hoff 2021) recruiting in the u.s. might result in cross-linguistic transference among participant responses. to control for this, the present study examined how mechanical support events were encoded in monolingual spanish-speaking children (4-to-6 year-old) and adults from a midsize urban latin-american city in colombia, where the national language is spanish. manizales is the capital of the department of caldas, and the population was around 450,000 people at the time of testing. the city is nestled between rural agricultural coffee-cultivating land and an urban environment with several renowned universities, contributing to the diversity in ways of life among its citizens. thirty-eight 4-to-6 year-old children (mage = 5 yrs, 9 mos.; range = 4 yrs, 0 mos. 6 yrs, 11 mos.; 16 females) were recruited from private preschool centers or elementary schools by sending home study flyers and consent forms. children were tested and recorded during school hours in a separate classroom or the school library (n = 36) or via zoom (n = 2). primary caregivers (n = 35; three caregivers did not report this information) self-reported education level based on a 6-point scale representative of standard educational trajectories in colombia at the time of testing (0 = n/a, 1 = preescolar [preschool], 2 = primaria [elementary], 3 = secundaria [high school], 4 = media [tech/trade school], 5 = universitaria [college-level], 6 = posgrado [graduate level]). most caregivers reported a post high-school equivalent level of education or higher (n = 27). thirty-two adults were recruited and tended to be caregivers of participating children, preschool or elementary school staff, or referred to by other participating adults. adults were tested in person (n = 10) or via zoom (n = 22) after school hours, and all responses were video recorded. proceedings of elm 3: 32-42, 2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta: mechanistic support language in colombian spanish-speakers. 35 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ten additional participants were tested, but their data were excluded due to not understanding the task/not completing the experiment (n = 6 children), suspected language delay (teacher-reported; n = 2 children), and video/connection error (n = 1 child, 1 adult). 3. procedure. participants were video recorded as they completed an elicited production task with a spanish-speaking and manizales-born research assistant in one-on-one sessions. the experiment began with a practice trial designed to familiarize participants with the procedure. during the practice, participants watched a motion event (a person rolling) and were asked to describe ‘what happened?’. all participants were shown sixteen videos (12 test trials, 4 filler trials) of dynamic support events. the videos were 7 to 9 seconds long and displayed a female agent acting out different dynamic configurations with cutout paper shapes (see figure 1). in the test trials, a figure object (paper) was attached to another object (tree or door) with a mechanism (clip, tape, or pin). all the configurations were designed such that they could be described with any of the verb types included in table 1. out of the test trials, half of the events showed the mechanism of attachment (i.e., visible) and the other 6 events concealed the mechanism of support (i.e., hidden). filler trials consisted of the female agent placing the paper cutout underneath an object (e.g., chair) or into another object (e.g., pot). after watching each trial, the researcher asked, ‘can you tell me what my sister did with my toy?’ in spanish. videos were replayed if requested by the participant. 4. coding. a trained research assistant transcribed and coded participant utterances (n = 816) for the verb type and subclass (see table 1). in addition, three other categories included other (verbs that did not fall into one of the four verb subclasses), no verb (utterances in which a verb was omitted), and does not encode support (utterances in which mechanical support was not described). we also coded for how the mechanism was encoded (see table 2). a second trained research assistant coded 14% of the transcriptions in terms of the verb type and subclass and how the mechanism was encoded. the inter-rater reliability was 93% and 97%, respectively. any disagreements between raters regarding coding were resolved by the final author. verb type verb subclass example in spanish non-mechanism simple verbs (blc events) poner = ‘put’ orientation verbs colgar = ‘hang’ mechanism general verbs of attaching pegar = ‘stick’ specific verbs of attaching enganchar = ‘hook’ table 1: verb classes relevant for encoding mechanical support (based on levin 1993) encoding of mechanism type example in spanish main verb lo pegó a la puerta = ‘she stuck it to the door’ prepositional phrase lo colgó con un gancho = ‘she hung it with a clip’ proceedings of elm 3: 32-42, 2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta: mechanistic support language in colombian spanish-speakers. 36 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ separate phrase le puso un clip y lo pusó ahí = ‘she put a clip on it and put it there’ main verb + prepositional phrase lo pegó con una cinta en una puerta = ‘she stuck it with tape on the door’ main verb + separate phrase le puso cinta negra y lo pegó en el armario = ‘she put black tape on it and stuck it on a dresser’ other (the mechanism was mentioned but not in terms of mechanical support between the figure and ground object) le puso un gancho = ‘she put a clip on it’ table 2: categories for encoding of mechanism 5. results. descriptive statistics for all verbs are reported in table 3. mechanism and nonmechanism verbs accounted for the majority of utterances (96.3%). other verb categories were less than 5%, so, the analyses excluded these categories. as shown in table 3, children used primarily simple verbs (e.g., poner = ‘put’), orientation verbs (colgar = ‘hang’), and general verbs of attaching (e.g., pegar = ‘stick’), whereas adults used primarily the latter two verb types (orientation verbs and general verbs of attaching). children adults visible hidden visible hidden non-mechanism simple verbs .28 (.03) .24 (.03) .09 (.02) .12 (.02) orientation verbs .24 (.03) .16 (.02) .41 (.04) .27 (.03) mechanism general verbs of attaching .44 (.03) .57 (.03) .40 (.04) .59 (.04) specific verbs of attaching .01 (.01) .03 (.01) other other .04 (.01) .006 (.01) no verb .004 (.004) .004 (.004) does not encode support .022 (.01) .027 (.01) .031 (.01) .017 (.01) table 3: mean proportions (with standard errors) of verbs for children and adults by visible and hidden stimuli first, focusing on the children, we conducted a mixed-effects logistic regression to examine whether children’s likelihood of using mechanism verbs (0 = non-mechanism verbs, 1 = mechanism verbs) increased with age between four and six years and whether mechanism visibility had an impact. age (in days) was entered as a covariate and mechanism visibility (visible vs. hidden) was entered as a fixed effect. random intercepts for participants, a by-participant random slope for mechanism visibility, and the correlation between the two were also entered. age significantly predicted the likelihood of using mechanism verbs; the use of mechanism verbs proceedings of elm 3: 32-42, 2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta: mechanistic support language in colombian spanish-speakers. 37 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ increased with age, (β = .7807, 95% confidence interval [ci] = 1.116 to 4.27, p = .023). however, mechanism visibility did not influence the likelihood of using mechanism verbs; children were equally likely to use mechanism verbs for events in which the mechanism was visible or hidden (β = .786, 95% ci = 0.663 to 7.26, p = .198). next, we explored the likelihood of using mechanism verbs (0 = non-mechanism verbs, 1 = mechanism verbs) among both children and adults (see figure 2). we conducted a mixed-effects logistic regression with age (children vs. adults) and mechanism visibility (visible vs. hidden) as fixed effects, and random intercepts for participants, a by-participant random slope for mechanism visibility, and the correlation between the two. age was not a significant predictor of verb type; children and adults did not differ in their use of mechanism verbs (β = .133, 95% ci = 0.525 to 2.487, p = .737). however, mechanism visibility significantly influenced the likelihood of using mechanism verbs; participants were more likely to use mechanism verbs when the mechanism was hidden compared to when it was visible (β = .808, 95% ci = 1.422 to 3.54, p < .001). figure 2: percentage of verb type (and subtype) included in utterances additionally, when participants used non-mechanism verbs (e.g., poner, colgar), we examined whether there were any differences in verb subclasses by age and visibility. children used more simple verbs (e.g., poner = ‘put’), whereas adults used more orientation verbs (e.g., colgar = ‘hang’; β = 3.052, 95% ci = 3.131 to 142.82, p = .002). mechanism visibility did not influence whether participants would use simple vs. orientation verbs, (β = .043, 95% ci = 0.328 to 3.329, p = .942). we also examined whether participants encoded mechanisms other than in the verb. we found that 62% of all utterances encoded the mechanism in a variety of linguistic structures (see table 2 and figure 3). while participants mostly encoded the mechanism in the verb (e.g., pegar = ‘stick’), proceedings of elm 3: 32-42, 2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta: mechanistic support language in colombian spanish-speakers. 38 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ and more for hidden events (58.2%) vs. visible events (44.2%), they encoded the mechanism in prepositional phrases more frequently for visible events (24.7%) than for hidden events (8.9%). figure 3: percentage of trials that encoded the mechanism (and how it was encoded) 6. discussion. our findings suggest that when describing events of an agent attaching a figure object to a ground object (e.g., girl puts paper on a tree with tape), both child and adult monolingual spanish speakers primarily use both non-mechanism and mechanism verbs; the use of mechanism verbs (e.g., pegar = ‘stick’) increases from 4 to 6 years of age, but then remains stable (i.e., there was no difference between children’s and adult’s use of mechanism verbs). in addition, mechanism verbs (e.g., pegar = ‘stick’) were more likely to be used when the mechanism was hidden vs. visible a finding that can be explained by examining how the mechanism was encoded in linguistic structures in addition to the verb. when the mechanism was visible, children and adults were more likely to encode the mechanism in a prepositional phrase (e.g., lo colgó con un gancho = ‘she hung it with a clip’) for visible events vs. hidden events; for hidden events participants were more likely to encode the mechanism in a general verb of attaching only (lo pegó a la puerta = ‘she stuck it to the door’). these findings shed light on the development of mechanical support language in spanishspeaking children, the influence of context—specifically, visibility of mechanism—on language, as well as the lexicalization patterns for encoding physical support in spanish more generally. we consider each of these in turn. the increase in the use of mechanical support verbs (e.g., pegar = ‘stick’) in children 4 to 6 years of age reflects the development of these verbs in english speaking children. for example, when 2.5to 4.5-year-olds describe mechanical support relations, they tend to use the blc, be on in english (e.g., the picture is on the wall), rather than lexical verbs that older children and adults use (e.g., the picture sticks/taped to the wall). it is not until about 6 years of age that englishspeaking children felicitously use a variety of lexical verbs and prepositions to encode mechanical support. johannes et al. (2016) propose that one reason for this protracted development is that be on may block the use of lexical verbs for english-learning children. in fact, when children are presented with a forced-choice task between lexical verbs and be on, they tend to choose the lexical verb, suggesting sensitivity to verb meanings (lakusta et al. 2024). similar studies should proceedings of elm 3: 32-42, 2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta: mechanistic support language in colombian spanish-speakers. 39 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ be conducted with spanish-speaking children to examine whether, when presented with poner (‘put’) vs. pegar (‘stick’), children would select pegar (‘stick’), further suggesting sensitivity to verb meaning. other explanations, not mutually exclusive from the first, are that children (regardless of the language they speak) have to learn about how mechanisms work—they must learn about the different causal force dynamic relations between the different mechanisms and the objects that they support (e.g., tape, magnet, glue). once this knowledge is acquired, children may use these types of verbs more felicitously in production. adult input may also play a role; studies in our lab are examining how parent input may change as children get older. the finding that mechanism visibility in the events affected how both children and adults describe mechanical support suggests that context plays a role in how event components get mapped to language. when the mechanism was visible (a picture attached to a tree with a clip), participants tended to encode it in a prepositional phrase (e.g., lo colgó con un gancho = ‘she hung it with a clip’); when it was hidden, a lexical verb was used (e.g., general verb of attaching – pegar = ‘stick’). this suggests that both children and adults 1) encode the mechanism when it is visible, and they do so in a prepositional phrase (see more on this below), and 2) make inferences about how the figure object is adhering to the ground object (e.g., by sticking) when the mechanism is hidden. these findings extend research reporting that children reason about mechanisms when they are engaged in exploratory play and try to find the causal structure of an object or system (e.g., how does this toy work? why did this toy stop working?; muentener & bonawitz 2017, schulz 2012). future research should further explore in what other ways context may play a role in the linguistic encoding of mechanical support events, such as examining whether certain types of mechanisms may be easier for children to describe than others (e.g., tape vs. magnets) and whether the effectiveness of the mechanism may influence children’s acquisition of the linguistic structures that encode it (e.g., if objects only sometimes stay up when using adhesive materials, this may influence children to pay attention to the mechanism and thus encode it more often in language). our findings also contribute information about how the spanish language lexicalizes mechanisms in physical support events. note that spanish speakers—adults and children—rarely used specific verbs of attaching (e.g., enganchar = ‘hook’). rather they used both general verbs of attaching (e.g., pegar = ‘stick’) and orientation verbs (e.g., colgar = ‘hang’); children also used simple verbs (e.g., poner = ‘put’). this pattern of verb use can likely be explained by the number of these verb types in spanish—spanish seems to have fewer specific verbs of attaching compared to other languages, such as english (which has more than 50; levin 1993), and thus it may not be surprising that spanish speakers would rarely use specific verbs of attaching to describe mechanical support. what seems notable, however, is that spanish speakers encode physical support with both orientation verbs and general verbs of attaching, whereas recent findings—that used highly similar support events as the current study—suggest that englishspeaking adults overwhelmingly use general and specific verbs of attaching (stick, tape, etc.) rather than orientation verbs (hang). although further research is needed that directly compares how english and spanish speakers encode mechanical support, we speculate that the two languages may show different lexicalization patterns—reflecting patterns that have been shown for how english and spanish encode manner of motion events (talmy 1985). for manner of motion events (the girl walked up the hill), english and spanish differ in how manner and path are lexicalized in the verb phrase (e.g., talmy 1985) english—a satellite-framed language—tends to encode the manner and motion in the verb (e.g., walk), whereas spanish—a path-framed language—tends to encode the path and motion in the verb (e.g., subir = ‘go up’). we propose that proceedings of elm 3: 32-42, 2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta: mechanistic support language in colombian spanish-speakers. 40 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ this lexicalization pattern may extend to the linguistic encoding of mechanical support events. mechanisms are intuitively similar to manners; how something is supported (mechanism) is akin to how something moves (manner). in contrast, the orientation of the figure relative to the ground is a component of the path (jackendoff 1992). given that mechanisms of support are types of manner (and spatial orientations are a component of the path), we predict that, in a study directly comparing the two languages, the pattern of language-specific lexicalization patterns observed for manner of motion events may extend to mechanical support events. if so, adult english speakers should lexicalize mechanical manner (i.e., mechanisms) in the verb phrase (over the orientation)— that is, they should primarily use mechanism verbs to encode mechanical support—whereas adult spanish speakers should do the opposite—they should use non-mechanism verbs to encode mechanical support and encode the mechanism in another clause. current research in our lab is testing this prediction. in sum, the current study sheds light on how children learn to talk about mechanical support events, and it does so by testing children and adults in colombia who are monolingual spanish speakers—thus diversifying the populations that are studied. the findings revealed that although both children and adults used non-mechanism (e.g., poner = ‘put’, colgar = ‘hang’) and mechanism verbs (e.g., pegar = ‘stick’) to describe events of a person attaching cutout paper shapes to another object, the use of mechanism verbs increased from 4 to 6 years of age. in addition, the visibility of the mechanism influenced how it was mapped to language; when the mechanism was visible (vs. when it was hidden), children and adults were more likely to encode the mechanism in a prepositional phrase (e.g., lo colgó con un gancho = ‘she hung it with a clip’). future research should explore the factors that explain the developmental progression that we observed in children and investigate other event features that may influence how language is mapped to mechanical support. additionally, it is important to examine how the encoding of mechanical support in spanish may compare to other languages, such as english, where adults primarily use mechanism verbs to encode the support (hauss et al. under revision). these findings shed light on the development of mechanical support language in spanish-speaking children, the influence of context—specifically, visibility of mechanism—on language, as well as the lexicalization patterns for encoding physical support in spanish more generally. references coventry, kenny r, richard carmichael & simon garrod. 1994. spatial prepositions, objectspecific function, and task requirements. journal of semantics 11. 289–311. https://doi.org/10.1093/jos/11.4.289 hauss, julia, jennifer barbosa, paul muentener & laura lakusta. the language of mechanical support in children: is it “sticking,” “hanging,” or simply “on”? [manuscript under revision]. herskovits, annette. 1986. language and spatial cognition: an interdisciplinary study of prepositions in english. cambridge: cambridge university press. jackendoff, ray s. 1992. semantic structures. cambridge: the mit press. johannes, kristen, colin wilson & barbara landau. 2016. the importance of lexical verbs in the acquisition of spatial prepositions: the case of in and on. cognition 157. 174–189. https://doi.org/10.1016/j.cognition.2016.08.022 johnston, judith r. & dan i. slobin. 1979. the development of locative expression in english, italian, serbo-croatian and turkish. journal of child language 6. 529-545. https://doi.org/10.1017/s030500090000252x proceedings of elm 3: 32-42, 2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta: mechanistic support language in colombian spanish-speakers. 41 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ lakusta, laura, julia wefferling, karima elgamal & barbara landau. 2024. “dividing the labor”: lexical verbs and the linguistic encoding of physical support in 2to 4.5-year-old children. journal of experimental child psychology 238. https://doi.org/10.1016/j.jecp.2023.105803 landau, barbara. 2020. learning simple spatial terms: core and more. topics in cognitive science 12(1), 91–114. https://doi.org/10.1111/tops.12394 levin, beth. 1993. english verb classes and alternations: a preliminary investigation. chicago, il: university of chicago press. levinson, stephen c. & david p. wilkins (eds.) 2006. grammars of space: explorations in cognitive diversity. cambridge: cambridge university press. https://doi.org/10.1017/cbo9780511486753 muentener, paul & elizabeth bonawitz. 2017. the development of causal reasoning. in: michael r. waldmann (ed.), the oxford handbook of causal reasoning, 677–698. new york, ny: oxford university press. https://doi.org/10.1093/oxfordhb/9780199399550.013.40 schulz, laura. 2012. the origins of inquiry: inductive inference and exploration in early childhood. trends in cognitive sciences 16. 382–389. https://doi.org/10.1016/j.tics.2012.06.004 talmy, leonard. 1985. lexicalization patterns: semantic structure in lexical forms. in: timothy shopen (ed.), language typology and syntactic description, 36-149. cambridge: cambridge university press. talmy, leonard. 2000. toward a cognitive semantics, volume 1: concept structuring systems. cambridge: the mit press. https://doi.org/10.7551/mitpress/6847.001.0001 vandeloise, claude. 1991. spatial prepositions: a case study from french. chicago, il: university of chicago press. welsh, stephanie n. & erika hoff. 2021. language exposure outside the home becomes more english-dominant from 30 to 60 months for children from spanish-speaking homes in the united states. the international journal of bilingualism: cross-disciplinary, crosslinguistic studies of language behavior 25(3). 483–499. https://doi.org/10.1177/1367006920951870 zimmermann, laura, lindsey foster, roberta m. golinkoff & kathy hirsh-pasek. 2018. spatial thinking and stem: how playing with blocks supports early math. american educator 42(4), 22-27. proceedings of elm 3: 32-42, 2025 jennifer barbosa, paola pinzón-henao, angelina pasquella, paul muentener, and laura lakusta: mechanistic support language in colombian spanish-speakers. 42 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ studying the interplay of context and semantic content in the interpretation of adversative conjunctions with eyetracking ghyslain cantin-savoie, denis foucambert, & grégoire winterstein* abstract. the french adversative connective mais, much like its english counterpart but, takes two conjuncts and indicates that they stand in some kind of opposition. the nature of this opposition is often discussed in the existing literature of adversatives. using an eyetracking experiment, we look at where and when this opposition appears to be manifested in a sentence reading task, using superiority comparatives sentences that appear degraded without a supportive context, as well as inferiority comparatives that do not seem to require such specific contextual information. we presented participants (n = 28) with four types of sentences, differing in the form of comparative, and in the information provided by the context. some contexts contained a pivot property to help readers access the opposition of the two conjuncts, while others were neutral in that regard. eye movements were recorded with a 250hz tobii pro fusion eyetracker, and mixed effect models were used to analyze the following eyetracking metrics : total fixation time and regression probability for the whole sentences, as well as first-fixation duration, first-gaze duration, regression-path duration, regressions-out and regressions-in for each individual words. even if the helping context is shown to lower negative acceptability judgment on sentences with mais plus (winterstein et al. 2014), we found no context effect in online reading processing, finding instead a persisting effect of the less/more dichotomy in all chosen measures, sometimes before fixation of those words, pointing towards a parafoveal effect of the aforementionned dichotomy, which merits closer look in future work. keywords. semantics; psycholinguistics; eyetracking; reading; adversatives 1. introduction. this work is centred on the semantics of the french adversative connective mais, typically rendered in english as but. the semantics of adversative connectives has been the topic of much work, in which two main theoretical branches can be distinguished. those branches can be characterized by the use of adversatives that they treat as canonical (winterstein 2017). on one hand, the contrastive approach considers that so-called contrastive uses of connectives like but exemplify best the core constraint indicated by adversatives (see a.o. umbach 2005). what seems to be required by but is that its conjuncts involve predicates that are comparable, i.e. similar in some respect and dissimilar in another. (1) contrastive john is tall but richard is small. on the other hand, inferential approaches take argumentative examples like (2) to be more representative of the semantics of adversatives (see a.o. blakemore 2002, winterstein 2012). crucially, *the authors would like to thank the audience of elm3 for useful comments. this work was supported by a frqsc soutien à la recherche pour la relève professorale grant (2022-np-296699). authors: ghyslain cantinsavoie, université du québec à montréak (cantin-savoie.ghyslain@courrier.uqam.ca), denis foucambert, université du québec à montréal (foucambert.denis@uqam.ca) & grégoire winterstein, université du québec à montréal (winterstein.gregoire@uqam.ca). proceedings of elm 3: 75-87, 2025 c©2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein published by the lsa with permission of the author(s) under a cc by license. 75 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the interpretation of such examples involves some additional inference, called a pivot, that connects the content of both conjuncts to a third proposition. those two connections must be dissimilar in some ways. (2) argumentative he’s unattractive but he’s kind so yes, i will marry him. in section 2, we discuss these approaches in more detail, highlighting a point of contention between them : the external or internal locus of the dissimilitude effect necessary for felicitous use of adversatives. we also show in that section that comparatives like the one in (3) allow us to control whether the required dissimilitude is internal to the conjuncts’ semantics (like it’s the case with less tall) or if an external proposition is needed to introduce it (like it’s the case for taller). (3) john is tall but less tall/*taller than richard. we then go on to describe a previous study by winterstein et al. (2014) who found that acceptability judgment discrepancy between negative (less tall) or positive (taller) comparisons can be alleviated with the introduction of an external proposition, whilst reading-time differences between positive and negative comparisons persist. this reading-time difference is what inspired our present eyetracking study, where we look for an online reading effect of the positive/negative comparison dichotomy (that we will call valence from now on) and where we see if this effect persists through the introduction of a preceding context made to help (or not) attaining required dissimilitude between conjuncts with an external proposition. our hypothesis is that there is a valence effect and that there is no effect of context type on reading time. in section 3, we describe the item-making, participant recruitment, and material-installing processes we went into to design an eyetracking experiment with the goal of testing the aforementioned hypothesis. filters, data transformation, eyetracking measures and statistical models are described in section 4. the results of those analysis are shown and discussed in 5. 2. empirical background: adversatives and comparatives. 2.1. contrastive and argumentative cases. let us start with a closer examination of examples (1) and (2). in the contrastive case of (1), john and richard refer to distinct individuals, who possess mutually exclusive properties (i.e. being small and being tall). in (2), the argumentative case oppose two properties that are not mutually exclusive (i.e. being unattractive and being kind) predicated of a single individual (the referent of he). though the examples differ in many ways, the one difference relevant to us is the locus of opposition, i.e. the question of determining from where stems the opposition that is required to license the use of an adversative like but as connective. for the contrastive use, the locus of opposition is internal to the conjuncts, i.e. supposed to be supported by the lexical semantics of the material in the conjuncts. this means that for a contrastive sentence with but to be felicitous, the second conjunct needs to deny a stated or implied proposition that is brought by the first conjunct itself. in the sentence in (1), john is tall brings an answer to the question who is tall, therefore implying that individuals, like john, can be in the set of tall people. it also implies the alternative possibility that other people, like richard, are also part of the tall people set. this said eventuality is denied by the second conjunct, richard is small, which proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 76 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ means that the use of but in this context is felicitous (see umbach 2005 for detailed account along those lines, crucially exploiting qud information to account for such cases). for the argumentative use of but, the locus of opposition is necessarily external, which means it will be found in the way both conjuncts are linked to a third, contextually given, proposition. as an example, in (4), he’s unattractive seems to go against i will marry him (indicated with a red arrow), while he’s kind does not (indicated with a green arrow). (4) buthe’s unattractive he’s kind i will marry him even if he’s unattractive and he’s kind are not necessarily opposed, it is through the opposing kind of connections to an external proposition that felicitous use of but is licensed. 2.2. comparatives cases. in this work we do not directly engage with the question of determining which of the two uses of adversatives just described is the most canonical (see winterstein 2012 for a discussion). we will however assume that some pivot inference is involved in the interpretation of adversatives, even in contrastive cases. to see why, we now introduce the core cases of this work, which we refer to as comparative cases, are exemplified in (5). (5) a. john is tall, but less tall than richard. b. #john is tall, but taller than richard. comparative cases appear similar to contrastive ones, given how both involve the comparison of two individuals along some property, like in (5). however, the property needs to be gradable, which is not a necessity for contrastive uses. tall being a gradable adjective, truth-evaluation needs a comparison of some kind. so, when it is said that john is tall, it inherently implies that john is taller than some threshold (kennedy & mcnally 2005). following the method described in 2.1, john is tall answer the question ”who is less tall than john?” by telling us that someone is in the set of things that are less tall than john ; if there is no one who is less tall than john, it cannot be said that john is tall. many people could possibly be in that set, including richard. the second conjunct of (5-a), [john is] less tall than richard denies the inclusion of richard in the aforementioned set, predicting a correct use of but. since [john is] taller than richard does not deny this implication, an internal-opposition-locus view of comparative uses of but seems to correctly predict the degraded nature of (5-b). when showed sentences like those in (5) in french, people judged sentences like (5-b) to be infelicitous (winterstein et al. 2014). reading times for sentences like those in (5-b) were also shown to be longer compared to (5-a). in the same experiment, another group of participants saw the same sentences presented after a context that introduces an external opposition pivot in the form of an equality condition, like in (6-a). (6) a. context: we are looking for a stunt double for richard, an actor, in action scenes for a movie. the stunt double needs to be exactly the same height as richard so that it goes unnoticed. the search is hard, because richard is tall. proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 77 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ b. sentence: john is tall, but less tall/taller than richard. c. opposition: butjohn is tall less tall/taller than richard john is the stunt double we are looking for when preceded by contexts like (6-a), sentences like (5-b) were not judged to be infelicitous anymore. however, reading time differences persisted, suggesting that the presentation of a context was not enough to facilitate the interpretation of the adversative connective. this suggests that those examples are somehow on a par with argumentative cases in that their interpretation rests on the identification of a contextual pivot. the reading times data also suggest that this identification comes at a later stage of interpretation. based on results obtained with the experiment of winterstein et al. (2014), we hypothesize that if we show sentences like (5) to participants, reading times will be greater on sentences with plus (more) like in (5-b) than the reading times on sentences with moins (less) like in (5-a). durationbased eye-tracking measures, described below in 4, should be sufficient to deny or confirm this hypothesis. secondly, since we think that the locus of opposition is internal, lack of opposition internal to the conjuncts cause by positive comparison with more should create a problem in online processes, regardless of context manipulations.1 therefore, our second hypothesis is that if we preceded sentences like those in (5) with helping contexts like in (6-a), the reading time effect of the less/more dichotomy will persist. to test this, we created another type of context, called ”neutral”, that are similar to the one in (6-a) but does not introduce an equality condition that allows an external opposition of conjuncts. if we are right, we should not see an effect of the nature of context on the eyetracking measures we observe. 3. method. to test these hypotheses, we ran an experiment in which participants had to read comparative sentences with mais preceded by contexts introducing (or not) an equality condition. the eye movements of the participants were recorded using an eyetracking device. this section describe the sentences that were shown, the hardware used, participants of the experiments, and the experimental protocol. 3.1. experimental items. each item had 3 parts : (1) a context, (2) a target sentence, and (3) a question. there were 60 items in total, of which 40 were distractors, meaning that the sentences were not comparatives with mais, so that subjects would not easily detect the objective of the experiment. 10 items and 30 distractors were extracted from winterstein et al. (2014) and slightly modified to be adapted to quebec french, the rest of the items were created for this experiment. the experiment used a 2x2 design, with the two binary conditions as follows. first, contexts from experimental items came in two versions (condition context). one version introduced an equality condition like the one shown in (6-a), and was called helping. the other version presented information irrelevant to the interpretation of the adversative, and is referred to as the neutral version. we illustrate these version in (7). 1this predict that for uses where we think opposition is external like argumentative cases in (2), context manipulation should have reading time effects. another experiment will be needed to look into this. proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 78 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (7) jeanne doe est amnésique et ne se rappelle plus qui elle est. les enquêteurs pensent qu’elle est la soeur jumelle/ petite soeur de marc leclerc, un vieil avocat porté disparu il y a 10 ans de cela. pour s’en assurer, ils font passer à jeanne une batterie de tests. translation : jane doe is amnesiac and doesn’t remember who she is. investigators think that she is the twin / little sister of marc leclerc, an old lawyer that disappeared 10 years ago. to be certain, they make jane go through a battery of tests. each context was modified or created so that the total number of characters in the context would be inside a 20 character range, between 222 and 242 characters. second, target sentences in the experimental items were comparative uses of mais. each sentence existed in two versions (condition valence), depending on the kind of comparative being used: either a superiority comparative (plus x que / ’more x than’) or inferiority one (moins x que / ’more x than’). crucially, each version only differs from the other in the choice of the comparative adverb. these two versions are shown in (8). (8) selon les tests, jeanne est vieille mais moins/plus vieille que marc même s’ils ont tous les deux les cheveux blancs. translation : according to the tests, jane is old, but less old/older than marc even if they both have white hair. the target sentences differ from previously shown examples in (5) in that they are surrounded by complement phrases, so that critical zones of interest (i.e. around the adversative connective) are not shown either at the extreme left or the extreme right of the screen when read. character number for the target sentences varied between 108 and 128. questions were introduced as a way to distract participants into thinking the answer was the criterion we evaluate, and to determine whether participants were really reading sentences or only pressing space while moving their eyes. questions were between 48 and 68 characters, each one asking for a yes/no answer that indicated comprehension of the previously read context-sentence pair, like shown in (9) about the sentence in (8) and the context in (7) (9) est-il possible que jeanne doe soit la soeur jumelle de marc leclerc? translation : is it possible that jane doe is the twin sister of marc leclerc? every part of an item, context, sentence or question, had a mean lexical frequency maximum of 737.84 , evaluated as the maximum threshold of the interquartile range. lexical frequency for each word was found using lexique3 database trained on the frantext corpus, which contains 218 french novels from 1950 to 2000 (new et al. 2004). 3.2. equipment. subjects sat on an adjustable desk-chair so that their heads, regardless of their individual height, would be at the same spot. their chin rested on a table model chin-rest, as they were looking at an asus tuf gaming vg259q computer screen of size 24.5” with a colour depth of 8-bit, resolution of 1920 per 1080 and a refresh rate of 59.94hz, placed 60 centimetres away from their face. their eye movements were observed with a tobii pro fusion 250 hz eyetracker. 3.3. participants. this article analyses the data of 28 participants, 13 male and 15 female, aged between 19 and 68, with a mean age of 31.5 years old. we invited them by going in their university class and sending facebook calls to student groups and friends. participants were compensated 15 proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 79 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ cad for their participation. participants were selected to have only quebec french as their first language, and to have a normal or corrected vision. their highest completed levels of education ranged from high school to phd. 3.4. procedure. after arriving and signing a consent form, participants were placed in front of a screen, rested their chins on the chin-rest, underwent calibration, for which they had to follow a moving dot displayed on the screen. they then completed a first test item after which they were given the chance to ask questions before proceeding with the rest of the experimental items. they were warned that before each item part, a fixation cross would appear and they would be asked to look directly at it before pressing the space bar to go to the next stimulus. they were asked to read each item part before pressing space, except after the question part, where they were asked to answer yes by pressing a or no by pressing l. to instantiate the experiment, we used the tobii pro lab software, which unfortunately does not offer latin-square style pseudo-randomization. we thus created four different trial groups, to ensure that contexts and sentences where shown so that every participants would see the same number of each condition, and each group would not see the same items in the same conditions than other groups, as shown in table 1. experimental items group 1 group 2 group 3 group 4 1-5 helping|less neutral|less helping|more neutral|more 6-10 neutral|more helping|less neutral|less helping|more 11-15 helping|more neutral|more helping|less neutral|less 16-20 neutral|less helping|more neutral|more helping|less table 1: condition repartition amongst participant groups after showing group 1 items to a participant, the next participant would see group 2 items, and so on to keep an even spread. after the experiment but before briefing, they would go through a reading span test in order to evaluate their respective working memory2. 4. analysis. since 28 participants read 20 sentences each, we would expect 560 sentence-level data points. however, item number 19 being one character too long, it showed on 2 lines on the screen and was therefore removed. our eyetracking device recorded valid gaze location data for 466 of the 532 remaining sentences. fixation data was then retrieved using the tobii-pro i-vt fixation filter (tobiiab 2023). 5 sentences were removed for containing less than 4 fixations each, resulting in a total of 461 sentence tokens. after removing fixation data on 10 words that were fixated less than 60 ms each, we had fixation data for a total of 5703 word tokens. data analysis was carried out on two separate levels : sentences and words. 4.1. the sentence level. at the sentence level, we took total fixation duration and regression probability as indicators of global sentence-reading difficulty (goldberg & kotval 1999). total fixation duration is defined as the sum, in ms, of the durations of fixations that happened in the reading of each sentence-part of the item triads. regression probability is the number of regressive saccades made during the recording of a sentence, divided by its total saccade count. each of those 2we don’t go in details in this article about this test’s procedure, it is described in marcotte (2014). proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 80 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ measures were took as dependent variables (dv) in their own mixed effect model generated with lmer() (bates et al. 2015) in r, sharing the formula shown below in (10). (10) dv ∼ valence + context + scolcal + span + (1|participant) + (1|media) valence is a categorical variable with two possible values : plus (’more’) or moins (’less’), that represents whether the word following mais (’but’) in item sentences was plus (i.e. a superiority comparative) or moins (i.e. an inferiority comparative). our hypothesis is directly linked to this variable. context is a categorical variable with two possible values, helping or neutral, that indicates whether the context preceding the sentence contains a helping opposition pivot in the form of an equality condition or not. our hypothesis is also directly linked to that variable. scolcal is an ordinal variable whose values range from 1 to 4. it is used to take into account the highest school level of participants, 1 for high school, 2 for cegep 3, 3 for a bachelor’s degree from university and 4 for masters, phds, and higher. span is a continuous variable representing the result of the reading span test, to control for working memory, which can affect several reading behaviours (traxler et al. 2012), as well as other potential individual cognitive abilities related to reading (kim et al. 2021). participant and media are factors with levels representing participant id and sentence number 4, respectively. the way they are instantiated in the model (1|variable) means that the intercepts, but not the slopes, are random, to allow a middle ground between model complexity and statistical power. they are interpreted by the model as cross-random effects giving us only 1 intraclass correlation coefficient (icc) per result table, which is acceptable since all of our participants saw all types of sentences. note that in the model presented in (10) we did not account for interactions between independent variables. models testing for interactions between our two hypothesis-making independent variables were compared to models without interactions. since there were no significant differences between the two in terms of fitting, we chose the simpler version to avoid problems linked to model complexity. 4.2. the word level. at the word level, we considered five different measures. for each word that was fixated at least once, we capture the duration of the first fixation (in ms). first fixation duration is linked to very fast first order cognitive processes like word and letter recognition (rayner 1998). we also included first gaze duration, which is the sum of all fixation durations in milliseconds on a given word from the first fixation until the subject fixates on something else. 87% of the time, according to our own data, first fixation duration and first gaze duration have the same value, since many words are only fixated once before a saccade brings the gaze on another one (rayner 1998). first-gaze duration are said to be influenced by cognitive processes that are not as fast as those influencing first fixation duration (rayner 1998), and also by plausibility of words in sentence context, which is not the case for first fixation duration (inhoff 1984). 3cegep (collège d’enseignement général et professionnel, roughly translated to college of general and professional teaching) is a school level unique to québec with students entering it at around 17 or 18 years old, created to make higher education accessible and affordable to a wider population. 4as an example, the sentence about jane doe previously shown is always number 13, regardless of whether it is in the plus or moins version. proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 81 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ fixation # jane is old but less old than marc 1 x 2 x 3 x 4 x 5 x 6 x 7 x table 2: fixation order of a hypothetical sentence reading regression-path duration for a word is the sum of all fixation durations, from the first one on the word until a fixation is made on another word to the right of it. in table 2, showing the ordering of fixation on a hypothetical sentence reading, the regression-path duration of the word but would be the sum duration of fixations 4, 5 and 6. regression-path duration is used to indicate processing difficulties oftentimes caused by parsing effects, such as syntactic ambiguities (hyönä et al. 2003). for each word, we also calculated its probability of being the starting or ending point of an ocular regression (regression-out, regressions-in). each of these measures were taken as the dependent variable of a lmer() model shown in (11), similar to the one in (10). (11) dv ∼ valence + context + scolcal + span + sfreqlex + nbcarac + (1|participant) + (1|media) the model in (11) and the choices leading to its form are identical to those described in section 4.1, with the exception of 2 additional fixed effects. nbcarac contains integer numbers representing word length by character number, which is said to influence if a word is going to be fixated on in the first place (rayner 1998). sfreqlex represents the lexical frequency of the word, taken from the lexique3 database made with the frantext corpus, transformed with the scale() r function, to avoid scaling issues brought by the large value range inherent to the frequency values. we analyzed the words in two steps. first, each word-level measure took turns being inserted as dependent variable in the model described above, allowing us to look for global effects of valence and context on words. secondly, we indexed words by position in the sentence, taking but as 0, the word before it as -1, the word after it as 1, and so on. we then selected a segment of the sentences we call critical zone going from indexes -3 to 4, where there is no variation in word type per index. as an example, the word before mais (but) at position -1 is always the first iteration of the gradable adjective. however, words at the position 7 can be articles, nouns or verbs without any given link between them, so taking position as a category would be a mistake. 5 5. results. 5some part-of-speech or semantic related tagging could alleviate this limit, and would be interesting in subsequent research. proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 82 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 5.1. sentence level: results. results at the sentence level in table 3 and table 4 show that sentences are fixated longer and that participants are more likely to look back when reading sentences with plus (’more’), which indicates that superiority comparatives increase sentence reading difficulty. however, context type does not affect those metrics, as found by winterstein et al. (2014). table 3: model results for the log transform of total fixation duration of sentences table 4: model results for regression probability of sentences sentence-level results, visualized in fig. 1 and fig. 2, confirm that valence has an effect on reading times. since our question is about where and when it happens, we move to results at the word level. figure 1: regression probability on sentences, split by experimental conditions figure 2: total fixation duration in milliseconds on sentences, split by experimental conditions proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 83 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 5.2. word level: results. at the word level, when we take all words into account, there is no significant effect of our variables on either first fixation or first gaze duration. however, if we test on individual word positions in the critical zone described in section 4.2, we find a significant effect of valence on first gaze duration on two word positions, shown on a white background on fig. 3. figure 3: first gaze duration per word position, white background = significant effect of valence the fact that first gaze duration is longer on those two words cannot be explained by saying that characters in the words or their lengths are harder to decode, since if it were the case there would also be a first fixation effect, which is not the case. also, and most importantly, the effect of the more/less dichotomy is seen on first gaze duration on the first occurrence of the gradable adjective (’tall’), which is before the occurrence of more/less. firstly, this indicates a limitation in metrics definition : we don’t know if the word more/less was fixated before tall in the cases where first gaze duration is higher. if it is not the case, it would indicate a parafoveal effect of valence on the reading of the word tall, which would need subsequent work to confirm. if we were to hypothesize an effect of context on a given metric, it would be on regression path duration, since it is said to represent not only the time spent decoding words, but also the time spent integrating visual level information with previously known words and facts stored in memory, such as personal knowledge or information given in a preceding context (cook & wei 2019). however, there was no effect of context on regression-path duration, not on all words taken as a whole nor on each individual word of the critical zone. there was, however, a significant valence effect on the individual word position for the second occurrence of the gradable adjective, shown on fig. 4 where the background is white. those results mean that when participants look at the second occurrence of the gradable adjective, they take more time and have to go back more in the superiority comparative cases. since we showed previously a significant valence effect on regression probability at the sentence level, it is no surprise that regressions-out and regressions-in probability on individual words also have a significant effect of valence when taken on all words. when we look at the proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 84 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 4: regression-path duration on word position, white background = significant effect of valence critical zone, we can see that there are more regressions coming from and out of the first occurrence of the gradable adjective, as shown in the white background of both fig. 5 and fig. 6. figure 5: regressions-out probability for each word position, white background indicating a significant effect of valence. figure 6: regressions-in probability for each word position, white background indicating a significant effect of valence. these results are not indicative of the fact that participants are regressing from the first occurrence of tall to itself, and is therefore indicative of a limitation of our study: there is no chronological information about saccade for now that allows us to know where the regressions from tall go, and where they come from. this would requires subsequent analysis. there is no context effect on regressions-in and regressions-out probabilities when taking all words together. however, looking at the critical zone allows us to see (fig. 7 and 8) that more regressions get out of the second named entity (pierre in the example) when the context is neutral to target the first named entity (paul examples). then again, subsequent analysis is needed to proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 85 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ discover if a significant number of saccades go from pierre to paul. figure 7: regressions-out probability for each word position, white background indicating a significant effect of context. figure 8: regressions-in probability for each word position, white background indicating a significant effect of context. 6. conclusion and outlooks. this study tested hypotheses regarding the reading time increase of comparative sentences with mais (but) comparing superiority and inferiority comparatives. theoretical work about the semantics of adversatives, as well as previous research using self-paced reading data (winterstein et al. 2014), led us to think that (1) sentences with plus (more) will have higher reading times, and that (2) contexts introducing (or not) an external proposition that explicitly supports the opposition between conjuncts should not affect reading times. results of an eyetracking study confirmed these hypotheses, additionally allowing us to point out some limits and future research avenues. sentences with a superiority comparative (plus ’more’) showed an increase in total fixation time for whole sentences as well as a localized increase of first gaze duration and regression-path duration, which seems to confirm our first hypothesis. although context type had no effect on duration-based metrics and therefore on reading times per say, confirming our second hypothesis, results from regressions patterns showed the influence of context on where participants decide to look back from and to. context type seems to influence whether participants will regress from and to named entities, whilst most of the localized duration effects were found on the gradable adjectives. limitations to our analysis were made apparent during this study, including lack of tagging on individual words that would have allowed us to categorize them by semantic function or syntactic category. further analysis of our data will allow us to inspect our finding that duration-based metrics show valence effects on the gradable adjectives while location-based metrics show valence effects on the named entities. a boundary paradigm study would be needed to confirm parafoveal effects influenced by the semantics of plus/moins. finally, another study with different items would be key to investigate whether different kind of adversative sentences (not only comparatives) with an external locus of opposition would have duration-based effects of context manipulation. proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 86 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ references bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67(1). 1–48. 10.18637/jss.v067.i01. blakemore, diane. 2002. relevance and linguistic meaning. the semantics and pragmatics of discourse markers. cambridge: cambridge university press. cook, anne e & wei wei. 2019. what can eye movements tell us about higher level comprehension? vision 3(3). 45. goldberg, joseph h & xerxes p kotval. 1999. computer interface evaluation using eye movements: methods and constructs. international journal of industrial ergonomics 24(6). 631– 645. hyönä, jukka, robert f lorch jr & mike rinck. 2003. eye movement measures to study global text processing. in the mind’s eye, 313–334. elsevier. inhoff, albrecht werner. 1984. two stages of word processing during eye fixations in the reading of prose. journal of verbal learning and verbal behavior 23(5). 612–624. kennedy, christopher & louise mcnally. 2005. scale structure, degree modification, and the semantics of gradable predicates. language 345–381. kim, young-suk grace, yaacov petscher & christian vorstius. 2021. the relations of online reading processes (eye movements) with working memory, emergent literacy skills, and reading proficiency. scientific studies of reading 25(4). 351–369. marcotte, sylvie. 2014. étude en temps réel de la révision de la morphographie du nombre du verbe chez les étudiants universitaires . new, boris, christophe pallier, marc brysbaert & ludovic ferrand. 2004. lexique 2: a new french lexical database. behavior research methods, instruments, & computers 36(3). 516– 524. rayner, keith. 1998. eye-movements in reading and information-processing: 20 years of research. psychological bulletin 124(3). 372–422. tobiiab. 2023. tobii pro lab. computer software. http://www.tobii.com/. traxler, matthew j, debra l long, kristen m tooley, clinton l johns, megan zirnstein & eunike jonathan. 2012. individual differences in eye-movements during reading: working memory and speed-of-processing effects. journal of eye movement research 5(1). umbach, carla. 2005. contrast and information structure: a focus-based analysis of but. linguistics 43(1). 207–232. winterstein, grégoire. 2012. what but-sentences argue for: a modern argumentative analysis of but. lingua 122(15). 1864–1885. 10.1016/j.lingua.2012.09.014. winterstein, grégoire. 2017. perspectives on argumentation within language. theoretical, processing, computational and social aspects.: université paris diderot–paris 7 dissertation. (habilitation à diriger les recherches). winterstein, grégoire, emilia ellsiepen, jaques jayes & barbara hemforth. 2014. effects of context on the processing of adversative and comparative constructions. in cuny, . proceedings of elm 3: 75-87, 2025 ghyslain cantin-savoie, denis foucambert, and grégoire winterstein: the interplay of context and semantic content in the interpretation of adversative conjunctions. 87 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ a nonce investigation of a possible conjunctive default for disjunction adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, & lyn tieu* abstract. our study explores whether there is a conjunctive default in the interpretation of disjunction, focusing on romanian children’s and adults’ understanding of nonce functional words. we investigate how participants interpret novel connectors such as mo and mo...mo, which could theoretically correspond to ‘(both) a and b’, ‘(either) a or b’, or ‘a not b’ / ‘neither a nor b’. our results reveal that both adults and children overwhelmingly assign a conjunctive meaning to these nonce words. this suggests the existence of a conjunctive default in interpreting unknown operators linking two elements, which could explain why children have sometimes been found to interpret disjunctions as conjunctions in previous studies (singh et al. 2016, tieu et al. 2017, bleotu et al. 2023). in particular, we discuss how this conjunctive default may influence romanian children’s interpretation of complex disjunctions such as fie...fie, potentially explaining why they treat these constructions conjunctively. importantly, our findings also raise broader questions about why certain logical interpretations are favored over others, and whether frequency or cognitive simplicity can drive such biases. keywords. conjunction; conjunctive default; disjunction; nonce words 1. main contribution. the current study explores whether romanian-speaking children and adults have a default preference for conjunctive interpretations when encountering two items a and b linked by an unknown operator. we connect these results to an explanation of why children sometimes interpret disjunction as conjunction. specifically, we examine what meaning romanian children and adults ascribe to novel functional words such as the single connector mo and the complex connector mo...mo (involving reduplication of mo) in the structures a mo b and mo a mo b. based on the distributional properties of romanian, in particular, the distribution of simple coordination, simple disjunction, and simple negation linking two nominals, possible interpretations of a mo b include the following: *this research is supported by the project “the acquisition of disjunction in romanian” pn-iii-p1-1.1-te-20210547 (te 140 din 30/05/2022) led by a. bleotu. a. nicolae was supported by the dfg grant ni-1850/2-1, as well as the erc synergy grant 856421 (leibnizdream). l. tieu was supported by the social sciences and humanities research council of canada and the connaught fund. a. benz’ s work was partly funded by the “linguistic meaning and bayesian modelling” project within the leibniz collaborative excellence programme (pi anton benz, application number: k535/2023). we are grateful to the undergraduate students at the faculty of foreign languages, university of bucharest, for taking part in the experiments. we also thank our research assistants for helping with data collection, and the children from no.248 kindergarten, dreamland and licurici kindergarten. we are also grateful to the audiences at linguistic evidence 2024 (22-23 february 2024, university of potsdam) and experiments in linguistic meaning 3 (12-14 june 2024, university of pennsylvania) for their useful comments and suggestions. authors: adina camelia bleotu, university of bucharest (adina.bleotu@lls.unibuc.ro) & andreea nicolae, zas berlin (nicolae@leibniz-zas.de) & mara panaitescu, university of bucharest (mara.panaitescu@lls.unibuc.ro) & gabriela bı̂lbı̂ie, university of bucharest (gabriela.bilbiie@lls.unibuc.ro) & anton benz, zas berlin (benz@leibniz-zas.de) & lyn tieu, university of toronto / western sydney university (marcs institute for brain, behaviour and development) / macquarie university (lyn.tieu@utoronto.ca). proceedings of elm 3: 53-64, 2025 c©2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu published by the lsa with permission of the author(s) under a cc by license. 53 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ • conjunctive reading: a s, i b (‘a and b’) • disjunctive reading: a sau b (‘a or b’) • negative reading: a nu b (‘a not b’) based on the distributional properties of correlative coordination, complex disjunction, and correlative negation linking two nominals, possible interpretations of mo a mo b include the following: • conjunctive reading: s, i a s, i b ‘both a and b’ • disjunctive reading: sau a sau b ‘either a or b’ • negative reading: nici a nici b ‘neither a nor b’ interestingly, while various interpretations seem to be allowed distributionally, our findings suggest that, when participants are exposed to sequences containing a connective that is unknown to them, they tend to favor a conjunctive interpretation over a disjunctive or negative one; that is, the conjunctive interpretation has a privileged status compared to the other possible interpretations. 2. disjunction in child and adult language. disjunctive statements may receive multiple interpretations. for example, a sentence like x acted upon objects a or b may be interpreted inclusively (such that x acted upon one object and possibly both a and b), exclusively (such that x acted upon one object, not both), and even conjunctively (such that x acted upon both objects, not just one). interestingly, while the inclusive interpretation is available to both children and adults, the exclusive interpretation is preferred by adults for both simple and complex disjunctions (nicolae & sauerland 2016, nicolae et al. 2024), while it is relatively rare in children (though see sauerland & yatsushiro 2018 for evidence that german children can be exclusive). the conjunctive interpretation, on the other hand, is specific to child language (singh et al. 2016, tieu et al. 2017, bleotu et al. 2023), but is absent from adult language. table 1 illustrates the available interpretations of disjunction in child and adult language for a sentence such as (1). (1) the hen pushed the train or the boat. interpretation paraphrase adults children inclusive the hen pushed one and possibly both. ✓ ✓ exclusive the hen pushed only one, not both. ✓ ? conjunctive the hen pushed both, not just one. ✗ ✓ table 1: possible interpretations of the disjunctive sentence the hen pushed the train or the boat in adults and children the inclusive interpretation of disjunction can be explained as a logical, literal interpretation of disjunction (noveck 2001), while the exclusive interpretation can be derived via the negation of the stronger conjunctive alternative a and b (grice 1975, 1989). the more controversial interpretation to explain is children’s conjunctive interpretation of disjunction. for this, different accounts have been proposed, which derive the reading as: (i) an implicature involving recursive exhaustification (singh et al. 2016, tieu et al. 2017),1 (ii) a basic meaning of disjunction alongside 1according to singh et al. (2016), the conjunctive interpretation is derived through recursive exhaustification as follows: first, children enrich the simple disjunct alternatives (yielding the hen only pushed the train, the hen only proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 54 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ inclusivity, according to the ambiguity approach (sauerland & yatsushiro 2018), and (iii) a mere experimental artifact arising when disjunctive statements exhaustively mention all of the objects that the character interacts or may interact with (skordos et al. 2020, huang & crain 2020). of particular interest to us is the second possibility, namely that in child language, disjunction may initially (also) have a conjunctive meaning, possibly by virtue of a conjunctive default which children fall back on when having to interpret more complex items/structures (such as disjunctive ones). 3. disjunction in child romanian. in an adapted version of tieu et al. (2017), bleotu et al. (2023) and bleotu et al. (2024c) examined the interpretation of disjunction in child and adult romanian. in order to test whether the conjunctive interpretation of disjunction is a mere experimental artifact rather than a genuine semantic or pragmatic interpretation, the authors compared cases where the pictured agent (corresponding to the sentential subject) was surrounded by two objects and the disjunctive statements mentioned both, with cases where the agent was surrounded by four objects and the disjunctive statements mentioned only two of these (see figure 1). given the abundance of disjunctive markers in romanian, the authors tested four different markers. two involved variants of the simplex disjunction sau: (i) sau with neutral prosody, where there is no prosodic boundary after the first disjunct, (ii) sau with marked prosody, where each disjunct receives stress, similar to the pattern seen in complex disjunctions. the other markers were complex disjunctions: (iii) sau...sau, which is a reduplicated form of the simple sau, comparable to the japanese ka...ka and ka or the french ou...ou and ou, and (iv) fie...fie, which has no simplex counterpart, much like the french disjunctions soit...soit versus ou. figure 1: examples of 2-object and 4-object displays for the sentence the hen pushed the train or the boat, from bleotu et al. (2024c) the results revealed that romanian 5-year-olds showed a consistent tendency to interpret all forms of sau-based disjunctions inclusively. however, for the complex disjunction fie...fie, there was evidence of both conjunctive and inclusive readings. furthermore, while a significant decrease in the conjunctive interpretation of fie...fie was found in the experimental set-up that involved four rather than two objects, overall, the conjunctive interpretation of fie...fie did not fully go away, remaining an available interpretation for children. this led bleotu et al. (2024c) to conclude that pushed the boat); they then exhaustify the disjunctive sentence with respect to these pre-exhaustified alternatives. this effectively amounts to a conjunctive interpretation: the hen pushed the train or the boat, but it is false that the hen only pushed the train, and it is false that the hen only pushed the boat. it is worth noting that this recursive exhaustification mechanism has been independently invoked to account for the derivation of free choice inferences associated with modalized disjunctive statements, such as you may push the train or the boat, in both adults and children (see, among many others, kratzer & shimoyama 2002 and fox 2007, as well as chemla & bott 2014 and tieu et al. 2016 for experimental evidence). proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 55 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ children’s conjunctive interpretation of disjunction is not a mere experimental artifact but a genuine linguistic interpretation. the current study further extends this investigation, asking whether children’s conjunctive interpretation of disjunction might be due to a conjunctive semantic default. 4. nonce paradigms. the investigation relies on nonce words as a method to probe into children’s syntactic bootstrapping, that is, their ability to interpret words by relying on syntactic cues (gleitman 1990, brown 1957). for instance, by relying on distributional information, children are able to differentiate nonce nouns (do you see a sib? / do you see any sib?) from nonce verbs (what is sibbing?). berko (1958)’s wug test brought further evidence that children extend known morphology to novel words, such as plural forms (one wug vs. two wugs) and verbal morphology (he zibs). many subsequent experiments followed, including naigles (1990), syrett & lidz (2010), yuan & fisher (2009), yuan et al. (2011, 2012), huang et al. (2021) among others, further supporting these findings. recent novel paradigms also investigate the existence of logical defaults in interpretation, for example, the human simulation paradigm (hsp; gillette et al. 1999), which tests whether adults can infer meaning from context (see dieuleveut et al. 2022 for application of this paradigm to modals) and artificial language learning paradigms (culbertson & schuler 2019, maldonado & culbertson 2021, 2022), which are used to investigate adults’ and children’s biases in learning artificial words. in this study, we will take the natural step of extending nonce paradigms to explore children’s and adults’ defaults in ascribing meaning to unknown logical operators. 5. experiments. 5.1. aim. in our study, we investigate the kinds of meanings children and adults ascribe to a sequence where two nouns are linked by nonce words. if there is a conjunctive default, we hypothesized that participants would default to interpreting the nonce connective as a conjunction. 5.2. procedure. we conducted a mo experiment, where a and b were linked by the nonce word mo (cf. a mo b), as well as a mo...mo experiment, where a and b were each preceded by the nonce word mo (cf. mo a...mo b), mimicking a complex connective. the two tasks we conducted employed a truth value judgment task in prediction mode (tieu et al. 2017) rather than description mode (singh et al. 2016), so as to license ignorance inferences, which often characterize disjunctive statements. participants were asked to evaluate whether a puppet named bibi correctly guessed the outcome of a situation. participants were told that bibi would sometimes make use of an unknown word, and they had to decide what it meant for bibi. they were also told that the unknown word does not refer to something that one can point to. participants had to say whether bibi guessed well. at the end of the experiments, participants had to say what they thought the nonce words meant. guesses took the form exemplified in (2) in the mo task and the form exemplified in (3) in the mo...mo task, and they were provided orally to participants. each trial involved three scenes, as shown in figure 2. (2) gǎina hen.def a has ı̂mpins pushed trenul train.def mo mo barca. boat.def ‘the hen pushed the train mo the boat.’ proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 56 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (3) gǎina hen.def a has ı̂mpins pushed mo mo trenul train.def mo mo barca. boat.def ‘the hen pushed mo the train mo the boat.’ scene 1 experimenter: there once was a hen who loved to play with her toys, and she especially loved to push them around! one day her papa gave her a train, a boat. the hen was very happy to play with them. let’s see if bibi can guess what happened next! scene 2 experimenter: bibi, tell us what happened next. bibi: the hen pushed the train mo the boat. (mo task) bibi: the hen pushed mo the train mo the boat. (mo...mo task) experimenter: let’s see if bibi’s right! scene 3 experimenter: look, the mouse carried this and this! so was bibi right? figure 2: the three scenes of an experimental trial in which the guess the hen pushed (mo) the train mo the boat was uttered in a 2-disjunct-true (2dt) context 5.3. materials. the test was preceded by two warm-up trials (one true, one false), consisting of a simple noun subject, a verb and a simple noun object, such as (4). the presence of warm-up items ensured that participants were familiarized with the procedure. these trials only contained words known to children, and crucially did not contain mo or mo...mo. (4) buburuza ladybug.def a has pictat painted cana. mug.def ‘the ladybug painted the mug.’ our test items involved four 1-disjunct-true (1dt) target trials (e.g., only the train was pushed), four 2-disjunct-true (2dt) target trials (e.g., both the train and the boat were pushed), and two 0disjunct-true (0dt) control trials (e.g., neither of the objects mentioned was pushed, but a different object was pushed).2 we also included three (true/false) fillers consisting of a simple noun subject, a verb, and a simple noun object, such as (5). (5) iepuras, ul bunny.def a has cules picked o a pară. pear ‘the bunny picked a pear.’ we avoided the use of logical operators such as conjunction, negation, or disjunction throughout 2as discussed in jasbi et al. (2018, 2022) and in recent work by bleotu et al. (2024b), children may have a tendency to produce more disjunctions and be more exclusive when a and b are incompatible with each other, e.g., the squirrel was either at the top or at the bottom of the tree, compared to when a and b are in principle mutually compatible. in this study, we restricted ourselves to situations where a and b are in principle mutually compatible, rather than mutually incompatible. proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 57 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the experiment so as to avoid priming participants in any way. in a previous study, bleotu et al. (2024a) showed that in the presence of relevant questions involving the conjunctive alternative (did the hen push the train and the boat?), children were more exclusive for sau-based disjunctions and more inclusive for fie...fie. based on these findings, we wanted to make sure that participants’ interpretation of the nonce words was not influenced by questions that included conjunctions or other logical operators. for the warm-up items and for the control items, the visual display showed the subject and three objects (only one of these was mentioned in the warm-up sentences, only two of these were mentioned in the 0dt disjunctive sentences). for the 1dt and 2dt targets, the visuals only included the subject doing the action and the two objects mentioned in the disjunctive utterances, with no additional objects pictured. 5.4. participants. 17 monolingual romanian-speaking children (3;06—5;11) and 21 romanian adult native speaker controls participated in the experiments. all participants first completed the mo task, followed by the mo...mo task two weeks later. 5.5. prediction. if there exists a conjunctive default for the interpretation of logical operators, we expect participants to interpret both mo and mo...mo conjunctively, in line with this default. 5.6. results. all participants passed the controls and fillers and were included in the analysis; overall accuracy was high on both fillers (95.15%) and controls (95.24%). we first analyzed the group data, focusing on responses to the 1dt condition. if participants showed a conjunctive preference, they should reject the 1dt targets, since only one of the disjuncts/conjuncts was verified; on the other hand, accepting 1dt targets would be consistent with either inclusive or exclusive interpretations. we fit a mixed effects logistic regression model in r (r core team 2021) to responses to the 1dt condition, with answer as a dependent variable (coded as 1 for yes, 0 for no), group (children vs. adults), disjunction type (mo vs. mo...mo) and their interaction as fixed effects, and random intercepts for participant and item. the model revealed a significant effect of disjunction (β = −0.89, se = 0.42, z = -2.1, p < .05) but no significant effect of group or interaction (both p > .05). next, we analyzed individual participants’ response patterns. based on their responses to the 1dt and 2dt targets, we categorized participants as: inclusive, exclusive, negative, conjunctive, or mixed. table 2 shows the expected pattern of responses for each category of interpretation in the mo task, while table 3 shows the expected response patterns in the mo...mo task. interpretation of ‘a mo b’ 1dt 2dt inclusive yes yes exclusive yes no negative yes (if a is true and b is false) no conjunctive no yes table 2: expected response patterns for 1dt and 2dt conditions in the mo task as shown in table 4, both children and adults preferred conjunctive interpretations of both mo and mo...mo. given the overall small numbers of participants, we conducted a fisher’s exact test to determine if children and adults differed in their distribution of interpretation types. we found proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 58 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ interpretation of ‘mo a mo b’ 1dt 2dt inclusive yes yes exclusive yes no negative no no conjunctive no yes table 3: expected response patterns for 1dt and 2dt conditions in the mo...mo task no difference between groups (p > .05); children and adults were equally conjunctive in their responses. we conducted an additional fisher’s exact test to compare the distribution of child responder types for mo and mo...mo: there were 12 conjunctive children and 5 non-conjunctive children in the mo task and 16 conjunctive children and 1 non-conjunctive child in the mo...mo task. the analysis revealed no significant association between the type of nonce operator and the observed number of conjunctive responders in children (p = 0.17, or = 0.16, 95% ci: 0.003 − 1.7). lastly, we also conducted the fisher’s exact test to compare responses for mo and mo...mo among adults: there were 13 conjunctive adults and 7 non-conjunctive adults for mo and 16 conjunctive adults and 4 non-conjunctive adults for mo...mo. the analysis revealed no significant association between the type of nonce operator and conjunctive responders in adults (p = 0.48, or = 0.47, 95% ci: 0.082− 2.4). group interpretation mo mo...mo children (n=17) conjunctive 12 16 negative 1 0 mixed 4 1 adults (n=20) conjunctive 13 16 negative 2 2 mixed 5 2 table 4: distribution of participants by interpretation in the mo and mo...mo tasks and 6. discussion. when adults and children are exposed to nonce words connecting a and b, their default interpretation seems to be conjunctive. even more strikingly, they seem to default to conjunction even in an experiment where bibi does not always make correct guesses, as evidenced by the fact that some of the fillers were true, while some were false. in the remainder of the paper we discuss some possible interpretations of the results: a processing approach, a frequency approach, a logical universal primitives approach, and variants of a strongest meaning preference approach. a processing approach according to a processing account, participants’ conjunctive preference could be due to a simplified processing of the a mo b and mo a mo b structures, leading them to disregard the unknown operators and interpret them as the mere juxtaposition of a and b. importantly, note that juxtaposing a and b leads to a conjunctive interpretation (winter 1995). thus, an utterance such as the hen pushed (mo) the train mo the boat may be understood as ‘the hen pushed the train, the boat’, and in turn as ‘the hen pushed the train and the boat’. under this approach, mo and mo...mo are essentially ignored. proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 59 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ there are, however, other approaches which assume that participants do not necessarily ignore the nonce operators but instead that they attribute meanings to them on the basis of frequency, of logical primitives, or due to a bias for strong meanings. a frequency approach one could consider a frequency-based approach, which takes into account the frequency of various interpretations (e.g., conjunctive/disjunctive). on such an approach, participants may simply associate the unknown connector(s) with the interpretation corresponding to the most frequent logical operator linking two elements, namely conjunction. such a view would be supported by corpus evidence from jasbi et al. (2018, 2022) that conjunction is more frequent than disjunction. a logical universal primitives approach yet another view would be to consider logical universal primitives to be more easily accessible than non-primitives, with both children and adults preferring to attribute to the nonce operators the meaning associated with a logical primitive. under a view which takes conjunction to be more basic than disjunction and conceptually simpler, the fact that both children and adults default to a conjunctive interpretation falls out. one such view is entertained by zimmermann (2000) and geurts (2005), according to whom disjunction is more complex and can be decomposed using conjunction. specifically, they claim that disjunctive interpretations can be analyzed as the conjunction of two modalized elements (♢a ∧ ♢b). strongest meaning preference according to a view which privileges strong meanings, it could be that participants opt for conjunction rather than disjunction because conjunction has the stronger meaning of the two (the hen pushed the train and the boat entails the hen pushed the train or the boat, but the reverse is not true). according to dalrymple et al. (1998), if a sentence is ambiguous between two meanings, people may prefer the stronger one. since sentences containing mo and mo...mo allow for multiple interpretations, we could assume participants observe this principle and choose the stronger meaning of conjunction over disjunction. in terms of acquisition, these findings are also in line with the subset principle (crain et al. 1994, crain & thornton 1998), a learnability principle which leads children to prefer stronger (subset) interpretations over weaker (superset) ones. starting off with stronger conjunctive meanings, children can learn the weaker disjunctive meanings via positive evidence. this would be preferable to a scenario in which children start off with weaker disjunctive meanings and then have to learn the stronger conjunctive meanings via negative evidence (which is known to be scarce). distinguishing between the accounts above is not straightforward, given that frequency may be a consequence of conjunction being a default, or a consequence of conjunction having a stronger meaning than disjunction. we consider it an important empirical finding that both children and adults seem to opt for conjunction over disjunction as the meaning of a nonce operator, and suggest that future research can attempt to disentangle the various possible explanations. our findings are also important in that they shed light on children’s conjunctive interpretations of disjunction. previous studies (see bleotu et al. 2023) have found that romanian children are conjunctive and inclusive in their interpretation of the complex disjunction fie...fie. while the inclusive interpretation could be explained as a preference for a logical/literal interpretation, one could instead argue that children’s conjunctive interpretation of fie...fie is due to a conjunctive default, especially if fie...fie is perceived as infrequent or less familiar to children (data from adult corpora suggest that fie...fie is less frequent than sau or sau...sau, see bleotu et al. 2023). proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 60 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ while our findings do not completely rule out the possibility that some children could also arrive at conjunctivity through an implicature, they do suggest that a likelier possibility is that most of the children rely on a conjunctive default when the operators they have to handle are unknown to them or more complex, which would be the case of disjunction, particularly for disjunctions such as fie...fie. on this story, the conjunctive meaning of disjunction would not be derived as an implicature but rather it would precede any implicature stage, which is also in line with recent work by aloni (2024). 7. conclusion. the nonce experiments presented in the current study suggest that both romanianspeaking children and adults show a conjunctive default in ascribing meaning to unknown operators linking two linguistic elements. thus, these findings support the idea that a conjunctive default could be a possible source for children’s conjunctive interpretation of disjunction. 8. future directions. further research can investigate whether the findings of this study are replicable cross-linguistically, by looking at children’s and adults’ behavior in languages other than romanian. such studies could investigate in what way differences in the distributional properties of conjunction and disjunction may affect the availability of a conjunctive default. moreover, it would be important to determine the contribution of the linguistic and visual components of the experiment. an outstanding question is whether our findings might actually reflect an experimental artifact, as argued by huang & crain (2020) and skordos & papafragou (2016). our experiment utilized two objects, both of which were explicitly mentioned in the sentences. one could wonder whether the observed preference for conjunction is affected by this particular visual display. to explore this further, future studies should replicate the experiment with four objects, allowing us to better assess the role of the visual component in shaping participants’ preferred interpretations. finally, further research is needed to investigate other possible sources for the conjunctive interpretation. for instance, we still do not know why, in child romanian, conjunctive meanings seem to arise mostly with the disjunction fie...fie, rather than with sau-based forms of disjunction. one possible account relies on the idea that there is syncretism between fie and the present subjunctive fie of the verb a fi ‘to be’ in romanian, possibly leading children to interpret fie...fie as ‘be it a, be it b’, and ultimately as ‘(there is) a and b’ by reducing the irrealis ‘be’ to a realis ‘is’ (bleotu et al. 2024c,d). future studies should further explore this possibility, as the conjunctive interpretation of disjunction could be the effect of various (joint) sources rather than attributable to a single source. it could also be that these sources play different roles at different stages of development, e.g., the conjunctive default could characterize the behavior of young children (say, 3and 4-year-olds), while errors of syncretism with the subjunctive could be at play in both younger and older children. a clear picture of the sources of the conjunctive interpretation of disjunction remains to be developed. references aloni, maria. 2024. neglect-zero and no-split: cognitive biases at the semantic-pragmatic interface. presentation at the workshop free choice inferences: theoretical and experimental approaches. berko, jean. 1958. the child’s learning of english morphology. word 14(2-3). 150–177. proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 61 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 10.1080/00437956.1958.11659661. https://doi.org/10.1080/00437956.1958. 11659661. bleotu, adina camelia, gabriela bı̂lbı̂ie, mara panaitescu, alexandre cremers, anton benz & lyn tieu. 2024a. does hearing and help children understand or? insights into scales and relevance from the acquisition of disjunction in child romanian. under review. bleotu, adina camelia, rodica ivan, andreea nicolae, gabriela bı̂lbı̂ie, anton benz, mara panaitescu & lyn tieu. 2023. not all complex disjunctions are alike: on inclusive and conjunctive interpretations in child romanian. proceedings of the annual conference of the cognitive science society 45. 3062–3069. bleotu, adina camelia, andreea nicolae, gabriela bı̂lbı̂ie, mara panaitescu, anton benz, andreea nicolae & lyn tieu. 2024b. the role of incompatible disjuncts in the acquisition of disjunction: insights from studies involving actual and missing logical words in child romanian. presentation at sinn und bedeutung 2024. bleotu, adina camelia, lyn tieu, anton benz, alexandre cremers, gabriela bı̂lbı̂ie, mara panaitescu & andreea nicolae. 2024c. children interpret some disjunctions conjunctively: evidence from child romanian. preprint on osf. https://doi.org/10.31234/osf. io/bywj2. bleotu, adina camelia, lyn tieu, gabriela bı̂lbı̂ie, anton benz, mara panaitescu & andreea nicolae. 2024d. on the conjunctive interpretation of the disjunction fie...fie in child romanian. proceeding of sinn und bedeutung 2023 . brown, r. w. 1957. linguistic determinism and the part of speech. the journal of abnormal and social psychology 55(1). 1–5. 10.1037/h0041199. https://doi.org/10.1037/ h0041199. chemla, emmanuel & lewis bott. 2014. processing inferences at the semantics/pragmatics frontier: disjunctions and free choice. cognition 130(3). 380–396. 10.1016/j.cognition.2013.11.013. crain, stephen, weijia ni & laura conway. 1994. learning, parsing and modularity. in charles clifton jr., lyn frazier & keith rayner (eds.), perspectives on sentence processing, 443–467. hillsdale, new jersey: lawrence erlbaum associates. crain, stephen & rosalind thornton. 1998. investigations in universal grammar: a guide to experiments on the acquisition of syntax and semantics. cambridge, ma: mit press. culbertson, jennifer & kathryn schuler. 2019. artificial language learning in children. annual review of linguistics 5. 353–373. 10.1146/annurev-linguistics-011718-012329. https: //doi.org/10.1146/annurev-linguistics-011718-012329. dalrymple, mary, makoto kanazawa, yookyung kim, sam mchombo & stanley peters. 1998. reciprocal expressions and the concept of reciprocity. linguistics and philosophy 21(2). 159–210. dieuleveut, anouk, annemarie van dooren, ailı́s cournane & valentine hacquard. 2022. finding the force: how children discern possibility and necessity modals. natural language semantics 30(3). 269–310. 10.1007/s11050-022-09196-4. fox, danny. 2007. free choice and the theory of scalar implicatures. in uli sauerland & penka stateva (eds.), presupposition and implicature in compositional semantics, 71–120. proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 62 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ basingstoke, uk: palgrave macmillan. geurts, bart. 2005. entertaining alternatives: disjunctions as modals. natural language semantics 13(4). 383–410. 10.1007/s11050-005-2052-4. gillette, jane, henry gleitman, lila gleitman & anne lederer. 1999. human simulations of vocabulary learning. cognition 73(2). 135–176. 10.1016/s0010-0277(99)00036-0. https: //doi.org/10.1016/s0010-0277(99)00036-0. gleitman, lila. 1990. the structural sources of verb meaning. language acquisition 1. 3–55. grice, paul. 1975. logic and conversation. in peter cole & jerry l. morgan (eds.), syntax and semantics, vol. 3, 41–58. new york: academic press. grice, paul. 1989. studies in the way of words. cambridge ma: harvard university press. huang, haiquan & stephen crain. 2020. when or is assigned a conjunctive inference in child language. language acquisition 27(1). 74–97. 10.1080/10489223.2019.1659273. huang, nick, aron white, chia-hsuan liao, valentine hacquard & jeffrey lidz. 2021. syntactic bootstrapping attitude verbs despite impoverished morphosyntax. language acquisition 29(1). 27–53. 10.1080/10489223.2021.1934686. https://doi.org/10.1080/ 10489223.2021.1934686. jasbi, masoud, akshay jaggi, eve v. clark & michael c. frank. 2022. context-dependent learning of linguistic disjunction. journal of child language 1–36. 10.1017/s0305000922000502. jasbi, masoud, akshay jaggi & michael c. frank. 2018. conceptual and prosodic cues in child-directed speech can help children learn the meaning of disjunction. cognitive science https://api.semanticscholar.org/corpusid:46676615. kratzer, angelika & junko shimoyama. 2002. indeterminate pronouns: the view from japanese. in yukio otsu (ed.), proceedings of the tokyo conference on psycholinguistics, vol. 3, 1–25. tokyo: hituzi syobo. maldonado, mora & jennifer culbertson. 2021. nobody doesn’t like negative concord. journal of psycholinguistic research 50(6). 1401–1416. paper=https://link.springer. com/article/10.1007/s10936-021-09816-w. maldonado, mora & jennifer culbertson. 2022. person of interest: experimental investigations into the learnability of person systems. linguistic inquiry 53(2). 295–336. paper=https://direct.mit.edu/ling/article-pdf/53/2/295/ 2014244/ling_a_00406.pdf. naigles, letitia r. 1990. children use syntax to learn verb meanings. journal of child language 17. 357–314. nicolae, andreea & uli sauerland. 2016. a contest of strength: or versus either–or. in polina berezovskaya nadine bade & anthea schöller (eds.), proceedings of sinn und bedeutung (sub), vol. 20, 551–568. open journal systems. nicolae, andreea c., aliona petrenco, anastasia tsilia & paul marty. 2024. exclusivity and exhaustivity of disjunction(s): a cross-linguistic study. to appear in proceedings of sinn und bedeutung (sub), vol. 28. noveck, ira. 2001. when children are more logical than adults. cognition 78. 165–188. r core team. 2021. r: a language and environment for statistical computing. r foundation for statistical computing vienna, austria. https://www.r-project.org/. proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 63 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ sauerland, uli & kazuko yatsushiro. 2018. the acquisition of disjunctions: evidence from german children. proceedings of sinn und bedeutung 21(2). 1065–1072. singh, raj, ken wexler, andrea astle-rahim, deepthi kamawar & danny fox. 2016. children interpret disjunction as conjunction: consequences for theories of implicature and child development. natural language semantics 24(4). 305–352. skordos, dimitrios, roman feiman, alan c. bale & david barner. 2020. do children interpret or conjunctively? journal of semantics 37(2). 247–267. skordos, dimitrios & anna papafragou. 2016. children’s derivation of scalar implicatures: alternatives and relevance. cognition 153. 6–18. syrett, kristen & jeffrey lidz. 2010. 30-month-olds use the distribution and meaning of adverbs to interpret novel adjectives. language learning and development 6(4). 258–282. 10.1080/15475440903507905. https://doi.org/10.1080/ 15475440903507905. tieu, lyn, jacopo romoli, peng zhou & stephen crain. 2016. children’s knowledge of free choice inferences and scalar implicatures. journal of semantics 33(2). 269–298. 10.1093/jos/ffv001. tieu, lyn, kazuko yatsushiro, alexandre cremers, jacopo romoli, uli sauerland & emmanuel chemla. 2017. on the role of alternatives in the acquisition of simple and complex disjunctions in french and japanese. journal of semantics 34(1). 127–152. winter, yoad. 1995. syncategorematic conjunction and structured meanings. in mandy simons & teresa galloway (eds.), proceedings of the 5th semantics and linguistic theory conference (salt), ithaca, ny: clc publications, cornell university. yuan, sylvia & cynthia fisher. 2009. “really? she blicked the baby?”: two-year-olds learn combinatorial facts about verbs by listening. psychological science 5. 619–626. yuan, sylvia, cynthia fisher, padmapriya kandhadai & anne fernald. 2011. you can stipe the pig and nerk the fork: learning to use verbs to predict nouns. in nick danis, kate mesh & hyunsuk sung (eds.), proceedings of the 35th annual boston university conference on language development (bucld), 665–677. boston, ma: cascadilla press. yuan, sylvia, cynthia fisher & jesse snedeker. 2012. counting the nouns: simple structural cues to verb meaning. child development 83. 1382–1399. zimmermann, thomas. 2000. free choice disjunction and epistemic possibility. natural language and semantics 8. 255–290. proceedings of elm 3: 53-64, 2025 adina camelia bleotu, andreea nicolae, mara panaitescu, gabriela bı̂lbı̂ie, anton benz, and lyn tieu: a nonce investigation of a possible conjunctive default for disjunction. 64 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ how to compute a focus: evidence from incremental processing morwenna hoeks, maziar toosarvandani & amanda rysling* abstract. the global interpretation of a focus marked sentence with a particle like only arises due to the interplay of several formal components: f-marking, the semantics of the particle, the nature of contrastive alternatives, and a dependence on context. in this paper, we argue that on-line reading measures can be used to probe which of these components are computed when. in order to isolate which components give rise to reading slowdowns typically observed on foci (birch & rayner 1997, benatar & clifton 2014, lowder & gordon 2015, hoeks et al. 2023), a maze reading study tested whether such slowdowns still arise on second-occurrence foci (sof)—foci whose inferences have already been computed in prior discourse and are therefore entirely predictable—and on foci whose size and location can only be determined via a previously introduced contrast. results indeed showed slowdowns on such foci, suggesting that these cannot solely be attributed to readers computing focal inferences anew, nor to comprehenders initiating their reasoning about the relevant alternatives. keywords. focus; alternatives; maze task; reading; second-occurrence focus; predictability 1. introduction. this paper is about the interpretation of strings like (1), which contain the focussensitive particle only and give rise to different exhaustivity inferences, depending on the size of its associated focus. (1) sarah only read a book about bats. a. sarah only read a book about [bats]f . ⇝ she did not read a book about anything else. b. sarah only read [a book about bats]f . ⇝ she did not read anything else. c. sarah only [read a book about bats]f . ⇝ she did not do anything else. we can think of the interpretation of such a string in at least two ways. the first is in terms of the inferences that it gives rise to from a global perspective, considering the entire sentence as a whole, as is done in theoretical semantics; the second is in terms of the string’s word-by-word interpretation, and how those inferences are deduced by a comprehender in real-time. the goal of this paper is to connect these two perspectives in a more direct way than has been done, by using incremental reading measures to probe, at a relatively fine-grained level, what aspects of a focus meaning is computed when. we present the results of a reading study, and argue that by investigating slowdowns that arise during reading of material in carefully controlled contexts, the incremental interpretation of focus can be linked up with the formal representations that describe its global meaning. these formal representations have several distinct components. therefore, what it means to understand the meaning of any of the sentences in (1a–c) is to know, minimally, the information *authors: morwenna hoeks, university of osnabrück (morwenna.hoeks@uni-osnabrueck.de), maziar toosarvandani, university of california santa cruz (mtoosarv@ucsc.edu) & amanda rysling, university of california santa cruz (rysling@ucsc.edu) proceedings of elm 3: 188-200, 2025 c©2025 morwenna hoeks, maziar toosarvandani, and amanda rysling published by the lsa with permission of the author(s) under a cc by license. 188 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ corresponding to each of these components. taking seriously the fact that the offline inferences triggered by focus arise due to an interplay between different pieces of formal machinery also suggests that the real-time interpretation of foci may be decomposable into several distinct subprocesses. below, we outline what these formal components are, before discussing how the incremental computation of their corresponding subprocesses can be probed using on-line measures. at the same time, we will use these subcomponents to structure a discussion of the preceding psycholinguistic literature, which has shown—using a variety of tasks and methods—that the presence of focus marking leads to a number of distinct behavioral effects. by conceptualizing the processing of focus in terms of these formal subcomponents, we can associate these behavioral effects with distinct subprocesses in real-time comprehension. while this will engender many questions, we will concentrate on the interaction between these subprocesses and the preceding discourse context. we will suggest that investigating the processing of second-occurrence focus (sof)—a focus for which some of the subcomponents have already been computed in prior discourse—uniquely sheds light on this issue, because it allows us to isolate those subcomponents responsible for structural predictions about focus from those subcomponents that give rise to the focus-related inferences and integration with the preceding context. in a maze reading study, we test reading of sentences with an sof in discourse contexts that specify the alternative sets before the interpretation of those foci. results from this experiment suggest that effects of focus still arise in reading measures when some aspects of a focus’ meaning have already been computed in prior discourse, and when accidental properties of focus are accounted for. 2. decomposing the computation of focus inferences. since natural language is not used in isolation, we consider (2), which adds a preceding discourse for the example in (1a). (2) lily read a book about whales and penguins, but sarah only read a book about [bats]f . ⇝ sarah didn’t read a book about whales or penguins. theoretical accounts of focus marking, designed to capture the exhaustivity inference arising in (2), typically involve formal representations that include at least four distinct components: (i) these accounts, first, include some abstract representation of focus, or f-marking as we will refer to it here, which is often understood as a feature that is part of the syntactic structure of a sentence (jackendoff 1972). in (2), for example, f-marking is placed solely on the constituent bats (indicated with a subscript f), and is ultimately responsible for both the phonological and interpretational effects of focus marking. (ii) the exhaustivity inferences arise because the presence of f-marking evokes a set of alternatives that contrast with the focus (jacobs 1983, rooth 1985, 1992) on bats, as in (2); the contrastive alternatives would be computed, roughly, by replacing bats within that sentence with other expressions that contrast with it. (iii) these evoked alternatives affect the basic meaning of a sentence as they interact with focussensitive expressions inside that sentence. for only, the evoked focus alternatives are negated, leading to an exhaustivity inference. (iv) finally, representations of the discourse context are also assumed to play a role in the interpretation of focus marked sentences: in rooth’s (1992) alternative semantics, for instance, proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 189 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ this effect is captured by requiring the alternatives evoked by a focus to be a subset of a set of alternatives made salient by the context. this contextual restriction leads, in the case of (2) specifically, to the inference that whales and penguins are among the contextually-relevant alternatives to the focus, and therefore among the things that sarah did not read a book about. in terms of a human comprehender interpreting the above sentence word-by-word in real time, there is a sense in which the pieces of information in (i-iv) may not be computed all at once. it could be, for instance, that in some cases the location of f-marking is determined prior to reasoning about the evoked alternatives, while in other cases the specification of an alternative set in the context may guide a reader’s assumptions about the exact location of an upcoming focus. this latter scenario may in fact be what happens when (2) is interpreted incrementally. as shown in (1a-c), the second clause of this sentence is in principle compatible with multiple focus structures. because a focal accent is only placed on the last word of this sentence, this string is ambiguous with respect to size of the f-marked constituent even in listening where comprehenders have full access to the prosodic signature. but for (2), the comprehender may still conclude that only bats is f-marked, because the preceding context sets up a contrast between whales, penguins and bats, signalling that the relevant alternative set consists of alternatives to those expressions (and not alternatives to larger constituents). here, information about the alternative set may thus be deduced from context before the f-marked phrase itself is encountered. assuming that interpreting focus may be decomposed into several non-contemporaneous subprocesses, as exemplified above, may also suggest that different aspects of incremental focus comprehension have their own distinct behavioral signatures, or that the exact nature of the effects observed on foci can depend on the information that is available before a focus is encountered. behaviorally, it has been shown that foci are more deeply encoded in memory than non-foci (singer 1976, mckoon et al. 1993, birch & garnsey 1995, gernsbacher & jescheniak 1995), and that focus marking leads to more accurate responses in changeand error-detection tasks (bredart & modolo 1988), as well as longer reading (birch & rayner 1997, benatar & clifton 2014, lowder & gordon 2015) and response times (hoeks et al. 2023). the comprehension of focused material has therefore been argued to require more processing resources than that of non-focused material. but in light of the above, the exact reasons such effects arise are not fully understood, as it still remains unclear what specific aspect(s) of focus comprehension are driving these observed effects. hoeks et al. (2023), for instance, have suggested that the reading slowdowns observed on foci can at least in part be explained by comprehenders having to infer what the contextuallyrelevant alternative set is. contrasting the way in which target sentences like (3) are read in multiple different contexts, they showed that slowdowns on focused material (apple pie) are diminished in contexts that specify the relevant alternatives to a focus, as in (3b), compared to cases where no alternatives to the focus were mentioned in the preceding context, as in (3a). (3) a. context a: did sarah want apple pie for dessert? no alt b. context b: did sarah want apple pie for dessert, or chocolate cake? alt target: sarah said it was [apple pie]f that she wanted for dessert in other words, readers spend more time interpreting a focus in the absence of such contextually specified alternatives, suggesting that perhaps part of the additional time it takes comprehenders to proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 190 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ read a focus is spent on inferring what the contextually relevant alternatives may be—component (iv) above. however, since their studies also found suggestive evidence that focus slowdowns do not entirely disappear in the presence of explicit contextual alternatives, it may be that other aspects of interpreting focus, like deducing the location of f-marking (i) or incorporating the alternatives with the sentence meaning (iii), also play a role. the observed reading slowdowns on foci are likely complex and may be driven by several aspects of focus comprehension at once. besides the additional inferences that need to be computed for focus, other (non-mutually exclusive) explanations for the observed focus slowdowns have been pointed out elsewhere in the psycholinguistic literature on focus. it has been suggested, for instance, that focus causes readers to slow down because foci are generally less predictable than non-foci, as they often constitute new information (benatar & clifton 2014). in general, readers may spend more time on material that is less predictable from their context, including some foci. but although discourse-new foci were indeed found to be read more slowly than foci that had already been mentioned in the preceding context, hoeks et al. (2023) also found that focus slowdowns still arose for foci that were themselves discourse-given, like the ones in (3). this suggests that these slowdowns can only in part be attributed to differences in newness/givenness. it may still be, however, that focus slows down reading because foci are unpredictable in some other way. for instance, even though apple pie had already been mentioned in both contexts preceding (3), this was also the phrase that was explicitly asked about in those preceding questions. the answer—i.e., what it is that sarah wanted for dessert—had not been established in the common ground up to that point by either of those contexts. the focused phrase may therefore have been less predictable than other material in this sense, not because it had not been mentioned before, but perhaps because that phrase resolves a thus far open question-under-discussion (qud). it has also been proposed that (some aspect of) the interpretation of focus marked material may generally be prioritized by the comprehension system (cutler & fodor 1979, morris & folk 1998). it may be that readers slow down on focused material, not only because it takes additional effort and time to compute the relevant inferences, but also because comprehenders strategically allocate more attention and processing resources to interpret them. perhaps material put in focus generally carries more important information, exactly because such material often answers an open qud, as in (3). such an account could explain the observation that foci are better remembered and attended to as well (singer 1976, cutler & fodor 1979, bredart & modolo 1988, mckoon et al. 1993, birch & garnsey 1995, gernsbacher & jescheniak 1995). finally, it may be that the fact that foci typically receive contrastive accents also plays a role in these focus slowdowns. it may be, for instance, that the human perceptual system is particularly attuned to allocating attention to focal accents and that this affects the general processing of language. such an explanation could also account for the reading slowdowns on focus, because it may be that the implicit prosodic structure that is assigned during silent reading could be used to differentially allocate resources (breen 2014)—resulting in longer reading times on material that is predicted to be implicitly accented as well (see e.g., lowder & gordon 2015 for an account that would be consistent with this hypothesis). in short, the behavioral effects observed on most foci could be attributed to any or all of these potential aspects of focus comprehension: such effects may be due to readers not expecting a focus, proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 191 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ or to the content of the focus somehow being new or unpredictable, or providing novel information in some other way; they may be due to the fact that a reasoning process about the contrastive alternatives to the focus is set in motion, or to the fact that a context in which such alternatives are relevant has to be accommodated; they could also be due to the fact that an (implicit) accent was present on those foci, and/or to the fact that more processing resources are being allocated to the focus, once this focus has been identified. in the basic scenario that has been tested thus far in the psycholinguistic literature, focus marking is typically unambiguously signaled by a particle, cleft or by the auditory signal, sometimes even in the absence of any preceding context. in those cases, it is hard to distinguish between all the possible processes related to the interpretation of foci, because they may all be triggered at once, the moment the focus is recognized as such. the goal of the present experiment is to eliminate some of these explanations. what will be crucial in motivating the design of the experiment is the observation that some aspects of focus processing may be set in motion before the focus is encountered, as we saw for (2). unlike previous work, the present experiment involved sentences that would be ambiguous with respect to the focus structure they contained in isolation, but for which context disambiguated the location of f-marking. this allows us to disentangle effects of focus from effects of focus prosody because the location of f-marking can be manipulated via context while keeping the prosody of that sentence constant. it also allows us to test how the presence of contrastive expressions in the preceding context affects reading of a subsequent focus. specifically, we hypothesize that if such preceding contexts guide readers in assigning focus marking further downstream, reasoning about the alternative sets must already be set in motion as the reader is first incrementally proceeding through the sentence, meaning that, on the focus itself, no preceding context would have to be accommodated in order to interpret it. this predicts the resulting focus slowdowns to either disappear altogether (showing that focus slowdowns are solely due to the novel computation of such alternatives), or to be diminished (in which case remaining slowdowns must index a cost associated with other components of focus processing). finally, we also compare novel foci with foci that are fully predictable within their discourse context, in order to determine to what extent focus slowdowns are driven by their relative unpredictability. we further motivate the design of this experiment next. 3. using second-occurrence focus to decompose focus slowdowns. this experiment used sentences with second-occurrence foci (sof), as in (4b), to probe the incremental processing of foci. (4) a. sarah read a book about whales and penguins, and bob only read a book about [bats]f1 b. no, lilyf2 only read a book about [bats]f1 these sentences have several properties which make them well-suited for current purposes. first discussed by partee (1999), they have since risen to fame mainly because only the focus on lily is marked by a pitch accent, while the focus on bats that is associated with the particle only is not. it must still be that the associate of only is underlyingly f-marked, because it must evoke a set of alternatives that is operated on by this particle. after all, (4b) still entails that lily did not read books about alternative topics besides bats. most theoretical accounts of such foci explain this mismatch between interpretation and prosodic realization by making reference to the fact that these foci have typically been mentioned as regular foci before their second occurrence as a focus, as (4a) illustrates (selkirk 2008, rooth 2010, beaver & velleman 2011). these two properties of proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 192 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ sof (non-accenting and givenness) allow us to measure effects of focus marking that cannot be due to material being either new or marked with a pitch accent. in fact, unlike the foci tested by hoeks et al. (2023), sof do not provide novel answers to a preceding question, nor do they provide novel information in any other way. the utterance that (4b) responds to already states that someone only read a book about bats, and it therefore does not just make the entire phrase only read a book about bats given, it also makes it already entailed by the discourse context. sof examples like these are therefore a good test case to study whether previously observed slowdowns on foci are due to them being unpredictable given the preceding discourse, or whether they are due to other aspects of focus comprehension. as outlined above, perhaps readers slow down on foci not only because they are updating their representation of the discourse context with incoming information provided by the focus itself, but also because they have to compute what the contextually relevant alternatives to a focus are. crucially, for sof it is not just the case that the focus itself is entailed, but also that the evoked alternatives, as well as the the relevant focal inferences, have already been computed during the interpretation of the previous sentence. already in (4a) it is the case that the particle only operates over the set of alternatives, negating alternative statements about types of things that could have been read about instead of bats. if the behavioral effects observed on foci are due to the novel computation of such inferences, we would expect such effects to be diminished for sof. comparing reading times on new foci with those on sof gives us an indication of how much such slowdowns are driven by the novel computation of focal inferences as well. moreover, these sentences crucially contain multiple foci which each serve to signal a distinct contrast, and are each associated with their own operator. such sentences thus require there to be multiple distinct sets of salient contextual alternatives. complex cases like these could potentially reveal how comprehenders deal with multiple such alternative sets. again, a comparison between new foci and sof, in particular, may tell us whether comprehenders are able to encode and maintain multiple alternative sets during their incremental assignment of focus structure, or whether it is only the alternative sets that have not been computed before that drive the observed slowdowns. because sof are fully predictable themselves, do not carry a focal accent, and do not introduce any novel inferences, it may very well be that slowdowns on these foci entirely disappear. since we test the reading strings for which preceding contexts disambiguate f-marking, the alternatives that are evoked by these foci must also be fully known at the point in time where these foci are encountered. even if focus slowdowns are due to the general prioritization of focused material, it is conceivable that such slowdowns disappear for sof, because such foci may not constitute important information. in fact, the comprehension of sof may be de-prioritized, as they occur inside the background of another focus that may receive higher priority instead. crucially, however, results of this experiment in fact found a consistent slowdown, even for sof, thus ruling out any explanation for the general focus slowdown in terms of newness, unpredictability, calculation of novel inferences, or accenting. next, we describe the design of this experiment in more detail. 3.1. method. previous work has established that focus marking causes readers to slow down, and the goal of the experiment presented here was to further investigate the exact source of such focus slowdowns. the design of the experiment did this in several ways: because the string of words of the target sentence, on its own, would be ambiguous with respect to the assignment of proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 193 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ f-marking, this experiment first tested to what extent salient sets of alternatives in the preceding context can guide comprehenders in their assignment of focus on subsequent material. moreover, by comparing reading times on second-occurrence foci with that of new foci, the experiment tested whether such slowdowns still arise even if the material that is being read is already entailed by the context and is entirely recoverable and predictable from it. finally, it also aimed to test whether focus slowdowns still occur on material that would not be accented if pronounced out loud. 3.1.1. materials. every item constituted a dialogue between two speakers, speaker a and speaker b, where the utterance by speaker b was considered the target sentence and the utterance of speaker a served as the context against which the target sentence was interpreted. the target sentence, uttered by speaker b, always contained the focus particle only which was placed inside that sentence in a position that was ambiguous in terms of the associate it takes. the preceding context sentence always consisted of two clauses in which two alternatives were contrasted with each other, and the size of the focus in the target sentence was therefore manipulated by the size of these alternatives, such that in the narrow conditions only single nouns were contrasted while in the wide conditions complex noun phrases consisting of two nouns were contrasted with each other. orthogonal to this manipulation of focus size, the utterance of speaker a also manipulated focus type since it determined whether the focused phrase uttered by speaker b was either new or second-occurrence (sof). an example item in all four conditions is shown in (5). (5) a. speaker a: abby read a book about penguins and whales, and bob read a book about [gorillas]f speaker b: and lily only read a book about [bats]f narrow new b. speaker a: abby read an article about penguins and a report on whales, and bob read [an article about gorillas]f speaker b: and lily only read [a book about bats]f wide new c. speaker a: abby read a book about penguins and whales, and bob only read a book about [bats]f speaker b: no, [lily]f only read a book about [bats]f narrow sof d. speaker a: abby read an article about penguins and a report on whales, and bob only read [a book about bats]f speaker b: no, [lily]f only read [a book about bats]f wide sof in all conditions, the first object noun in the target sentence (book) constituted the critical region of interest, because—if focus structure in the context affected focus structure on the target sentence— it is this region that would either be focused (in the wide conditions) or not (in the narrow conditions). since the material inside the target sentence was held constant across conditions, response time differences in this region between the wide and narrow conditions could only indicate a difference in the projected focus structure of this sentence. if an effect of size can be observed on this first np, this means that readers indeed must be able to use previously specified contrasts to guide their assignment of focus marking. if a slowdown can be found in the wide sof relative to the narrow sof conditions on this region, this would indicate a slowdown due to focus marking, and would thus indicate that such slowdowns cannot solely be explained in terms of their unpredictability, accenting, or the accommodation of alternatives in the context. proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 194 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ in total, 48 items like (5) were constructed, each with the four conditions illustrated above. another 64 filler items which also consisted of multi-line discourses were interspersed with test stimuli. using a latin square design, all 48 items were counterbalanced over 4 lists, such that each participant saw one condition from every item. 3.1.2. procedure & participants. target sentences were presented using the maze task, while context sentences were presented normally. the maze task is similar to the more commonly used self-paced reading task in that response times are measured using button presses. but instead of simply pressing a button to advance to a following word each time a participant has read the current word, participants in the maze task see each word in the target sentence presented alongside a distractor word (or foil). participants must at every new word choose the correct continuation between the intended item and its foil, which would not make a sensical continuation. foils were automatically generated using the automaze software developed by boyce et al. (2020). an example of the automaze output for one target sentence is given in (6) below. on the second line, the distractor word is presented below its corresponding word of the target sentence. (6) no, x-x-x lily came only fine read call a ew book been about trump bats, must, but jack i hill might glass be laws misremembering hypothyroidism it. am. on every trial, participants first read a context sentence on one screen. on a subsequent screen, participants were presented with the start of the target sentence in the format of the maze task. that is, only the utterance of speaker b was presented incrementally; the utterance of speaker a was presented all at once for normal reading. the context sentence disappeared from the screen when participants moved on to the target sentence. to ensure careful reading of the context, all experimental trials were followed by a comprehension question that probed various properties of the context preceding the target sentence. for instance, the example item in (5) was followed by the comprehension question in (7). (7) did speaker a mention books about sharks? when participants chose the wrong maze word (the foil), they automatically exited the maze trial and were sent immediately to the comprehension question. before being presented with the target stimuli and fillers, participants read a short description of the task, followed by five practice items. practice items were similar to experimental items in that they involved a short context sentence, followed by a maze-sentence and a comprehension question. after the short practice phase, the experimental items were presented along with the fillers in a pseudo-random order. 55 native speakers of english were recruited via prolific and compensated at a $12 hourly rate. data from 48 participants were included in the analysis; 7 participants were excluded because they failed to complete more than 70% of the maze sentences. 3.2. results. the mean comprehension question accuracy was 79%, and the mean completion rate of the maze target sentences of experiment 1 was 89%. mean response times per condition for each region are given in table 1 and plotted with 95% confidence intervals in figure 1. data were analyzed using r, version 3.6.3 (r core team 2021). bayesian (generalized) linear mixed-effect models were fit using stan, as implemented in the brms package, version 2.18.0 (bürkner 2017), with the default priors. separate models were fit to log-transformed response proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 195 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ condition subject particle verb np1 np2 narrow new 834.10 (23.74) 803.18 (14.11) 696.93 (11.48) 789.48 (13.22) 954.33 (15.80) wide new 810.50 (17.55) 844.61 (17.10) 728.04 (16.33) 936.04 (17.26) 939.57 (20.26) narrow sof 971.72 (31.47) 753.85 (14.45) 704.20 (12.08) 774.20 (14.89) 862.86 (16.41) wide sof 947.68 (25.50) 771.36 (14.67) 714.24 (12.70) 859.30 (17.59) 886.34 (16.95) table 1: experiment 1: mean rt and standard error of the mean in each condition two words before, at, and two words after the target word. figure 1: experiment 1: mean rt in each region in each condition. error bars represent the 95% confidence interval. times and untransformed response times as dependent measures. models included populationlevel effects of focus and newness (coded 0.5, -0.5), with narrow focus and new conditions treated as reference levels, and random slopes and intercepts for both subjects and items (baayen et al. 2008). for each model, we ran four chains, each with 5000 steps (warmup = 1000 steps). rhat statistics in all models approached 1.00 and no warnings emerged. table 2 and 3 present the posterior estimates of the fixed effects as well as 95% credible intervals for both models of experiment 1. below, reliable effects on each region will be discussed. no reliable effects were found at either the verb region, or the spillover regions of np2. on the subject (lily), positive estimates for focus type indicate that subjects in the sof conditions were read reliably slower than subjects in the new conditions. contrastive focus on the sof subjects may thus have induced a slowdown relative to new conditions in which subjects were also new but were not contrastively focused. on (only), positive estimates for type indicate that this particle was read faster in the sof conditions than in the new focus conditions. this may be expected as this particle was new in the new but not in the sof conditions. on the first np (book), positive estimates for focus size indicate that this noun was read faster in the wide conditions than in the narrow conditions, and positive estimates for focus type indicate that sof foci were read faster than new foci. there was also a reliable interaction between focus size and type, though the focus size effect was reliable among both the sof (β =71.89, proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 196 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ est. error 95%cri rhat intcpt 2.90 .01 [2.87,2.92] 1.00 size 0.03 .01 [0.02,0.04] 1.00 type 0.05 .01 [0.03,0.07] 1.00 si:ty 0.04 .01 [0.01,0.06] 1.00 table 2: posterior estimates for the population-level effects of logrts on np1. est. error 95%cri rhat intcpt 884.00 38.94 [807, 962] 1.00 size 51.90 21.26 [9.50,94.1] 1.00 type 120.00 28.62 [64.1,175] 1.00 si:ty 96.22 44.93 [7.94,186] 1.00 table 3: posterior estimates for the population-level effects of raw rts on np1. 95%cri=[3.51, 140.59]) and the new conditions (β=168.11, 95%cri=[94.73,242.04]). positive estimates for size on the np1 spillover region again indicate that this phrase was read faster in the narrow than in the wide conditions. however, the estimated credible interval for the main effect of focus type as well as the interaction overlapped with zero so these will not be considered reliable. finally, on the second np (bats), positive estimates for focus type again indicate that np2 was read slower in the sof than in the new conditions, but estimates for size and the interaction were not reliable. this may be unsurprising because the second np was focused in both wide and narrow conditions, and therefore the only difference between the conditions was their newness. 4. discussion. by carefully manipulating the context preceding a focus, the experiment presented here was designed to disentangle various explanations for the reading slowdowns typically observed on foci, with the goal to better understand what aspects of a focus’ meaning are computed when. to test whether focus slowdowns still arise on second-occurrence foci and in cases where their alternative sets were already fully determined in the prior context, the experiment crossed the type of focus (new or sof) with the size of that focus inside a target sentence. reading times were indeed found to be affected by the manipulation in focus size: slowdowns were found on nps that were put in focus in the wide focus condition, relative to the narrow focus condition in which that np was not f-marked. slowdowns were thus found at the left edges of wide foci, indicating that the source of these slowdowns cannot lie in the anticipation of a focal accent, which would only be assigned at their right edge, if assigned at all. since the size of these foci could only be determined via a previously specified contrast, these results also confirmed that contextual alternatives can guide readers’ projection of f-marking onto subsequent material. the results of this experiment revealed an effect of focus size even in the sof conditions. this slowdown allows us to draw several conclusions about the processing of focus: it first indicates that focus slowdowns cannot be explained solely in terms of predictability, because these sof were fully recoverable from their context and did not present any new information at the point where they were encountered. it also indicates that slowdowns can be found even for foci that had already occurred as foci, and whose corresponding inferences had already been computed. the presence of such slowdowns, although diminished compared to focus slowdowns in the new conditions, suggests that these focus slowdowns cannot solely be attributed to a cost associated with computing focal inferences anew, nor can they be explained by comprehenders starting to set a process in motion that involves reasoning about these alternatives. after all, if readers had not proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 197 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ already encoded some information about these alternatives prior to the occurrence of the sof, they would not have been able to tell whether these phrases themselves were f-marked at all. one possibility is that the sof slowdown on np1 does not indicate the assignment of focus marking, but instead indicates that readers slowed down because they expected new material in those regions even though those regions were in fact always given. but even in that case, some explanation would be needed for the size effect: since the only difference between the narrow sof and wide sof conditions was the size of the alternatives that were mentioned in the preceding context, all else being equal, reading times in these conditions could only have been affected by this contextual manipulation. even if readers slowed down because the target sentence was in some way different than expected, such expectations would still have to have been formed based on the alternatives mentioned in the context. any explanation of this effect would need to make reference to this contextual manipulation of alternatives, which would ultimately entail that some representation involving these alternatives guides readers’ downstream behavior. there are, as far as we can tell, only three possible explanations for the sof slowdowns that remain. first, it may be readers slow down on such foci because computing their inferences is a relatively automatic process triggered by the presence of a particle signaling an upcoming focus— i.e., a process triggered even if the current sentence is entirely parallel to a preceding sentence for which such inferences have just been computed. alternatively, it could be that such slowdowns indicate additional care and attention for the processing of foci, which would have to hold, again, even for foci that have just been interpreted. note that such prioritization itself therefore cannot be explained in terms of a general prioritization of material that is new or provides a novel answer to an (implicit) question. finally, it may be that these sof slowdowns arose because there is some level of representation at which events and their participants are being linked together, and that the contrastive focus on a subject like lily results in an over-writing process where subsequent material describing a specific type of event is actively being linked to that new individual, after it has been de-linked from the previously encoded agent of that event (bob). however, whatever the exact characterization of this cost may be, it is clear from these results that the cost itself must be sensitive to the preceding discourse that sets up a particular alternative set: in all three cases, readers must still be sensitive to the particular size of the contrast involved. this suggests that comprehenders must somehow encode a representation of that discourse context which goes beyond the linear organization of the preceding sentences, e.g., involving the kind of representation adopted in a qud-based framework—one which would represent, for the examples involved in our experiment, an overarching qud asking either who read what or who read a book about what topic. finally, what these results moreover suggests, at a more general level, is that the interpretation of focus has consequences for the comprehension system that cannot be reduced to some accidental property of focus, like accenting, newness, or predictability. they suggest, instead, that the very presence of focus marking may be driving these effects, even if that focus does not contain any new information or introduce any novel inferences. references baayen, r. harald, douglas j. davidson & douglas m. bates. 2008. mixed-effects modeling with crossed random effects for subjects and items. journal of memory and language 59(4). proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 198 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 390–412. https://doi.org/10.1016/j.jml.2007.12.005. beaver, david & dan velleman. 2011. the communicative significance of primary and secondary accents. lingua 121(11). 1671–1692. benatar, ashley & charles clifton. 2014. newness, givenness and discourse updating: evidence from eye movements. journal of memory and language 71(1). 1–16. https://doi.org/10.1016/j.jml.2013.10.003. birch, stacy l. & susan m. garnsey. 1995. the effect of focus on memory for words in sentences. journal of memory and language 34(2). 232–267. https://doi.org/10.1006/jmla.1995.1011. birch, stacy l. & keith rayner. 1997. linguistic focus affects eye movements during reading. memory & cognition 25(5). 653–660. https://doi.org/10.3758/bf03211306. boyce, veronica, richard futrell & roger p. levy. 2020. maze made easy: better and easier measurement of incremental processing difficulty. journal of memory and language 111. 104082. https://doi.org/10.1016/j.jml.2019.104082. bredart, serge & karin modolo. 1988. moses strikes again: focalization effect on a semantic illusion. acta psychologica 67(2). 135–144. https://doi.org/10.1016/0001-6918(88)90009-1. breen, mara. 2014. empirical investigations of the role of implicit prosody in sentence processing. language and linguistics compass 8(2). 37–50. https://doi.org/10.1111/lnc3.12061. bürkner, paul-christian. 2017. brms: an r package for bayesian multilevel models using stan. journal of statistical software 80. 1–28. 10.18637/jss.v080.i01. cutler, anne & jerry a fodor. 1979. semantic focus and sentence comprehension. cognition 7(1). 49–59. https://doi.org/10.1016/0010-0277(79)90010-6. gernsbacher, morton ann & jörg d. jescheniak. 1995. cataphoric devices in spoken discourse. cognitive psychology 29(1). 24–58. https://doi.org/10.1006/cogp.1995.1011. hoeks, morwenna, maziar toosarvandani & amanda rysling. 2023. processing of linguistic focus depends on contrastive alternatives. journal of memory and language 132. 104444. https://doi.org/10.1016/j.jml.2023.104444. jackendoff, ray. 1972. semantic interpretation in generative grammar. cambr., ma: mit press. jacobs, joachim. 1983. fokus und skalen: zur syntax und semantik der gradpartikeln im heutigen deutsch. tübingen: niemeyer. lowder, matthew & peter gordon. 2015. focus takes time: structural effects on reading. psychonomic bulletin & review 22(6). 1733–1738. https://doi.org/10.3758/s13423-015-0843-2. mckoon, gail, roger ratcliff, gregory ward & richard sproat. 1993. syntactic prominence effects on discourse processes. journal of memory and language 32(5). 593–607. https://doi.org/10.1006/jmla.1993.1004. morris, robin k. & jocelyn r. folk. 1998. focus as a contextual priming mechanism in reading. memory & cognition 26(6). 1313–1322. https://doi.org/10.3758/bf03201203. partee, barbara h. 1999. focus, quantification, and semantics-pragmatics issues. in: peter bosch & rob van der sandt (eds.), focus: linguistic, cognitive, and computational perspectives, 213–231. cambridge: cambridge university press. r core team. 2021. r: a language and environment for statistical computing. rooth, mats. 1985. association with focus. amherst, ma: umass amherst dissertation. rooth, mats. 1992. a theory of focus interpretation. natural language semantics 1(1). 75–116. proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 199 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ https://doi.org/10.1007/bf02342617. rooth, mats. 2010. second occurrence focus and relativized stress f. in: caroline fery & malte zimmermann (eds.), information stucture: theoretical, typological, and experimental approaches, 15–35. cambridge: cambridge university press. selkirk, elisabeth. 2008. contrastive focus, givenness, and the unmarked status of “discoursenew”. acta linguistica hungarica 55(3–4). 1–16. singer, murray. 1976. thematic structure and the integration of linguistic information. journal of verbal learning and verbal behavior 15(5). 549–558. proceedings of elm 3: 188-200, 2025 morwenna hoeks, maziar toosarvandani, and amanda rysling: how to compute a focus: evidence from incremental processing. 200 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ pseudo-scoping out of tensed clauses: cumulation vs. buildups jonathan palucci* abstract. tensed complement clauses are often assumed to be scope islands for quantifier raising (qr) of universal quantifiers. however, as observed by farkas & giannakidou 1996, barker 2022, hoeks et al. 2022, a.o., there are apparent counterexamples to this assumption, where a universal dp appears to scope out of a tensed complement clause to take scope over a singular indefinite in the matrix clause, henceforth ‘variation readings’. hoeks et al. 2022 propose that qr out of tensed clauses is possible, but only in event structural configurations which involve buildup processes. in this paper, we report experimental results providing evidence that variation readings are not sensitive to buildups. we then offer an alternative analysis, capturing variation readings as a form of cumulation, and we present experimental results supporting this analysis. the empirical generalization suggests that tensed complement clauses are islands for qr after all, and apparent counterexamples are due to other mechanisms. keywords. quantifier raising; universal quantification; scope islands; tensed complement clauses; cumulativity; event structure; pseudo-scope 1. introducton. tensed complement clauses are often assumed to be scope islands for quantifier raising (qr) of universal quantifiers, as illustrated in (1) (chomsky 1975, may 1977). (1) a student claimed that every speaker had a ride. a. available: a single student claimed that all the speakers had a ride. b. unavailable: for every speaker x, a student claimed x had a ride the only available interpretation for the example in (1) is the one paraphrased in (1-a), where the universal quantifier scopes within the tensed complement clause of claim. if the universal dp, every speaker, were able to undergo qr out of the tensed complement, we would expect a reading where the universal quantifier scopes above the singular indefinite, as paraphrased in (1-b). however, as observed by farkas & giannakidou 1996, barker 2022, hoeks et al. 2022, a.o., there are counterexamples to this generalization. one such counterexample is provided in (2). (2) a student made sure that every speaker had a ride. a. available: a single student made sure that all the speakers had a ride. b. available: for every speaker x, a student made sure x had a ride. unlike (1), the example in (2) licenses a reading where the universal dp, every speaker, seemingly scopes above the singular indefinite, as paraphrased in (2-b)—henceforth, a ‘variation reading’. the contrast between (1) and (2) raises the following empirical puzzle: variation readings appear to be predicate sensitive (barker 2022). for example, the variation reading in (2-b) involves the lf in (3), where we posit an instance of ‘non-local’ qr out of the tensed complement clause. *thanks to michael wagner, bernhard schwarz, four anonymous elm reviewers and members of the syntaxsemantics reading group at mcgill for feedback and comments. this work is supported by a sshrc doctoral fellowship (752-2023-2295) and an frqsc doctoral scholarship (2022-b2z-306924). authors: jonathan palucci, mcgill university (jonathan.palucci@mail.mcgill.ca). proceedings of elm 3: 276-288, 2025 c©2025 jonathan palucci published by the lsa with permission of the author(s) under a cc by license. 276 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (3) lf: [ [ every speaker ] λ1 [ [ a student ] make sure tp[. . . t1 . . . ] ] ] if one were to capture variation readings through qr, this raises the corresponding theoretical puzzle: why is an lf like (3) available for (2) but not (1)? 2. buildup approach (hoeks et al. 2022). one possible response to the empirical/theoretical puzzle introduced above is to posit a restriction on non-local qr, on a predicate-by-predicate basis (barker 2022). the trouble is that this solution is rather stipulative. a more elegant response to this puzzle introduced above is presented in hoeks et al. 2022. the idea is that non-local qr is a mechanism that is, in principle, available in the grammar (anderson 2004, syrett 2015, wurmbrand 2018), but it is only possible in certain event structural configurations. in particular, they propose that the matrix event must involve a buildup process in order for non-local qr to be licensed. intuitively, a buildup process can be thought of as an eventuality where the quantification is ‘built up to’ over time by individual cases of the quantification. this proposal explains the contrast between (2) and (1) as follows. the lexical semantics of the predicate make sure involves such a buildup process, which is why non-local qr is possible in (2). in contrast, the lexical semantics of claim doesn’t involve such a buildup process, meaning that non-local qr is not available in (1). building on the contrast between (1) and (2), hoeks et al. 2022 make two main claims. the first claim concerns licensing variation readings with the embedding predicate claim: two manipulations to a sentence like (1) should make the variation reading available (henceforth, ‘external buildup cues’). the first manipulation involves adding an adverbial, like by 8pm, which signals that the matrix event is construed as a buildup process. the second manipulation involves changing the aspect of the embedding predicate to perfect aspect. the intuition behind this manipulation is that perfect aspect signals that the buildup process has lead to a result state. the second claim concerns the embedding predicates which can license variation readings. more specifically, external buildup cues only work for a subset of embedding predicates—predicates which they identify as ‘buildupicle’ predicates. examples of buildupicle predicates provided in hoeks et al. 2022 include: claim, heard, found, become aware, believe/come to believe. in contrast, they also provide examples of ‘non-buildupicle predicates’: is confident, is sure, is aware, is convinced, realize, remember. 3. experiment 1: testing the buildup approach. hoeks et al. 2022 make the following empirical claim: external buildup cues license a variation reading for (4), in contrast to (1). (4) by 8pm, a student had claimed that every professor had a ride. (hoeks et al. 2022; p.444) the sentence in (4) serves as a crucial data point as it provides the basis for our first experiment. in particular, we had two goals in our first experiment: i) to test whether the claim concerning (4) is borne out and ii) to test whether the availability of variation readings differed between buildupicle and non-buildupicle predicates. the buildup approach predicts that: i) a variation reading should be more readily available with (4), in contrast to (1); ii) in general, variation readings should be more readily available with buildupicle predicates, in contrast to non-buildupicle predicates. 3.1. experimental design. to test these claims, we conducted a sentence rating experiment with 20 participants. participants were recruited on prolific, and we restricted participation to those who self-report to be native speakers of north american english, having grown up and currently live in the united states or canada. we also restricted the experiment to participants with at proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 277 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ least a 97% approval rating on prolific. participants were shown context-sentence pairs and asked to rate how natural the sentence sounded (given the context) on a 6-point likert scale, where 1 corresponded to ‘completely unacceptable’ and 6 corresponded to ‘completely acceptable’. for the experimental design, we manipulated predicate type (buildupicle predicate vs. nonbuildupicle predicate) and context type (buildup context vs non-buildup context). there were a total of 12 item sets, where each item set involved a different embedding predicate. every participant saw every condition in each item set for a total of 48 trials. to minimize spill-over effects between sentences from the same item set, their distance was maximized by creating 4 blocks of stimuli, each with a latin square design so that a participant saw only one condition from each item set in that block and the same number of trials from each condition. in addition, order within each block was randomized, and the order between blocks was also randomized between participants. the list of of predicates is provided in (5). six of the predicates were buildupicle predicates, so the buildup approach predicts them to pattern like claim in terms of licensing variation readings with singular indefinite subjects in the presence of external buildup cues. the other six predicates were non-buildupicle predicates, so the buildup approach predicts them to not license variation readings with singular indefinite subjects, even in the presence of external buildup cues. (5) a. buildupicle predicates: claim, heard, found, become aware, believe/come to believe b. non-buildupicle predicates: is confident, is sure, is aware, is convinced, realize, remember in the experiment, each item set comprised of four conditions and the context varied in each condition. there were two variations of the target sentence: one involving no buildup cues and one involving buildup cues. two of the conditions involved the no buildup cue variant while the other two conditions involved the buildup cue variant. as a reminder, the buildup cue variant made use of two manipulations: i) the target sentence contained a buildup adverbial (like by the end of the talk) and ii) the embedding predicate contained perfect aspect (i.e., had claimed) in one condition (‘non-buildup context, non-varying’; i.e., the control condition), the context involved a single individual, so that the singular indefinite in the target sentence referred to a single individual (corresponding to a ‘surface scope’ reading). in this condition, neither the target sentence nor the context involved any external buildup cues, (6). in another condition (‘nonbuildup context, varying’), the context involved several individuals, so that the singular indefinite in the target sentence varied between individuals (corresponding to an ‘inverse scope’ reading). in this condition, neither the target sentence nor the context involved any external buildup cues, (7). in the third condition (‘buildup context, non-varying’), the context again involved a single individual, so that the singular indefinite in the target sentence referred to a single individual (corresponding to a ‘surface scope’ reading). however, in this condition, both the context and the target sentence involved external buildup cues, (8). in the fourth and final condition (‘buildup context, varying’), the context involved several individuals, so that the singular indefinite in the target sentence varied between individuals (corresponding to an ‘inverse scope’ reading). again, in this condition, both the context and the target sentence involved external buildup cues, (9). participants were given the chance to report any issues that arose throughout the experiment and none of the participants reported any issues. at the end of the experiment, participants were also given the opportunity to guess what they believed the experiment was about and none of the participants ascertained the proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 278 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ goal of the experiment (this is important since participants were exposed to many similar stimuli). (6) no buildup, non-varying indefinite context: [bea is a student. at last week’s talk, the speaker presented three theories in total. in the final discussion, bea raised issues with each theory and said they were all wrong.] a student claimed that every theory was wrong. (7) no buildup, varying indefinite context: [ann, bea and carol are students. at yesterday’s talk, the speaker presented three theories. during the final discussion, ann claimed the first theory was wrong, bea claimed the second theory was wrong and carol claimed the third theory was wrong.] a student claimed that every theory was wrong. (8) buildup, non-varying indefinite context: [ann is a student. during last week’s invited talk, the speaker presented three different theories in total. when the speaker presented the first theory, ann raised her hand and claimed the theory was wrong. then, when the speaker presented the second theory, ann raised her hand and claimed the theory was wrong. finally, when the speaker presented the third and final theory, ann again raised her hand and claimed the theory was wrong.] by the end of the talk, a student had claimed that every theory was wrong. (9) buildup, varying indefinite context: [ann, bea and carol are students. during yesterday’s talk, the speaker presented three theories in total. when the speaker presented the first theory, ann claimed it was wrong. when the speaker presented the second theory, bea claimed it was wrong. finally, when the speaker presented the third theory, carol claimed it was wrong.] by the end of the talk, a student had claimed that every theory was wrong. 3.2. results. the results for experiment 1 are in figure 1. the main takeaway is that the varying condition with buildup cues is rated significantly worse than the non-varying condition; and more importantly, the varying condition with buildup cues is rated just as bad as the varying condition without buildup cues. this suggests that the external buildup cues suggested by hoeks et al. 2022 didn’t facilitate the availability of variation readings. in addition, there doesn’t appear to be any difference between buildupicle and non-buildupicle predicates (in terms of licensing variation readings). the results were analyzed using linear mixed effects models (using the lme4 package in r) with predicate type and context as fixed effects—including interactions—and random intercept and slopes by item and participant (including interactions).1 the statistical model indicated a clear contrast between the varying and non-varying conditions overall, indicating that variation readings are either not present or at least much harder to get compared to the non-varying baseline. the factors predicted to facilitate variation readings by the buildup approach had no effect though: the statistical model indicates no significant difference between buildupipcle and non-buildupicle predicates (β = -0.0521, p = 0.68) and no significant interaction between predicate type and context (β = 0.1208, p = 0.62). these null results suggest that, if the distinction between buildupicle and non-buildupicale predicates and the presence of buildup cues have any effect at all, they are not 1model specification to analyze results: (i) lmer(response ∼ buildupicle.vs.not*varying.vs.not + (1+varying.vs.not || item) + (1+buildupicle.vs.not*varying.vs.not || participant) proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 279 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ large enough to be detectable with an experiment this size, suggesting that they do not play a major role in explaining when variation readings are available. buildup context non−buildup context non−varying varying non−varying varying 0 2 4 0 2 4 re sp on se predicate type buildupicle non−buildupicle acceptability ratings for variation readings of (non)−buildupicle predicates figure 1: left: non-varying and varying indefinite contexts involving buildups, (8)–(9). right: non-varying and varying indefinite contexts involving no buildup, (6)–(7). 4. when are variation readings available?. in this section, we propose an alternative analysis of variation readings. we argue that variation readings don’t involve non-local qr but arise from certain inferential properties of the embedding predicate (drawing on harada 2022). more specifically, we observe that embedding predicates which license variation readings with singular indefinites also license a form of predicate sensitive cumulation with plural subjects, suggesting a connection between the two phenomena. we refer to this as ‘the cumulating approach’ (palucci 2024). more generally, a goal in the remainder of the paper is to argue that mechanisms other than qr are needed to analyze variation readings; the cumulating approach is one possible analysis. to understand what we mean by ‘cumulation’, consider (10). the target sentence is true in the provided context under a reading that is weaker than the distributive reading one might expect. the cumulation we are concerned with reflects the fact that the source of distributivity, whatever it is, can be absent and the resulting truth conditions reflect certain inferential properties of the embedding predicate which allow us to combine ann and bea’s contributions together. thus, in (10), the two separate instances of ‘making sure’ from ann and bea can be cumulated together (due to the cumulating properties of make sure) so that ann and bea, between them, made sure that every problem was error-free. for this reason, the target sentence is felicitous in this context. we refer to these inferential properties, whatever they may be, as ‘cumulation’. (10) plural subject context: [ann and bea are teaching assistants. they were asked to review four homework problems. ann made sure the first and second problems were errorfree, but didn’t look at the third and fourth problems. bea made sure the third and fourth problems were error-free, but didn’t look at the first and second problems.] ann and bea made sure that every problem was error-free. the difference between make sure and claim then boils down to whether each predicate has the proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 280 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ necessary inferential properties. for example, consider (11). in this case, ann and bea’s claims can’t be combined into a single claim since claims from different individuals aren’t the kind of thing that can be cumulated. for this reason, the target sentence is not felicitous in this context. (11) plural subject context: [ann and bea are teaching assistants. they were asked to review four homework problems. ann claimed that the first and second problems contained errors, but had no issues with the other problems. bea claimed that the third and fourth problems contained errors, but had no issues with the other problems.] #ann and bea claimed that every problem contained errors. furthermore, we propose that the contrast between make sure and claim represents a more general divide between two classes of embedding predicates. on the one hand, we propose that there are predicates that pattern like make sure—henceforth, ‘cumulating predicates’. on the other hand, we propose that there are predicates that pattern like claim—henceforth, ‘non-cumulating predicates’. the hypothesis underlying the cumulating approach is that the availability of variation readings correlates with this kind of predicate sensitive cumulation; in other words, variation readings are more readily available with cumulating predicates, in contrast to non-cumulating predicates. also note that the cumulating approach makes a different prediction regarding the data point in (4). the crucial prediction by the buildup approach was that external buildup cues will make nonlocal qr available even when the universal dp is embedded in the tensed complement clause of claim. in contrast, the cumulating approach predicts (4) to be bad—regardless of external buildup cues (since claim doesn’t have the necessary inferential/cumulating properties). 4.1. experimental design. to test the above claims, we conducted 2 sentence rating experiments, each with 32 participants. participants were recruited on prolific., with the same restrictions as the first experiment. participants were shown context-sentence pairs and asked to rate how natural the sentence sounded (given the context) on a 6-point likert scale, where 1 corresponded to ‘completely unacceptable’ and 6 corresponded to ‘completely acceptable’. one experiment looked at variation readings with singular indefinite subjects, while the other experiment looked at cumulation with plural subjects. these experiments aimed to replicate the experiments in palucci 2024 but with an improved experimental design, better stimuli, more predicates and more participants. for the design of the experimental stimuli, we manipulated predicate type (cumulating predicate vs. non-cumulating predicate) and context type (varying context vs. non-varying context). for both experiments, there were a total of 20 item sets, where each item set involved a different embedding predicate. the list of predicates is provided in (12). ten of the predicates (cumulating predicates) were hypothesized to pattern like make sure in terms of licensing cumulation with plural subjects and variation readings with singular indefinite subjects. the other ten predicates (non-cumulating predicates) were hypothesized to pattern like claim in terms of not licensing cumulation with plural subjects and variation readings with singular indefinite subjects. (12) a. cumulating predicates: make sure, confirm, establish, prove, verify, determine, guarantee, corroborate, double-check, note down b. non-cumulating predicates: claim, notice, confess, heard, believe, hope, fear, know, say, realize proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 281 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ for the experiment on variation readings, each item set comprised of four conditions. the target sentence was the same in each condition (i.e., a singular indefinite subject, an embedding predicate and a tensed complement clause containing a universal quantifier) but the context varied. in one condition (i.e., the control condition), the context involved a single individual, so that the singular indefinite in the target sentence referred to a single individual (corresponding to a ‘surface scope’ reading), (13). in another condition, the context involved several individuals, so that the singular indefinite in the target sentence varied between individuals (corresponding to an ‘inverse scope’ reading), (14). the other two conditions were unacceptable baselines. in one baseline, the context again involved a single individual but the embedded clause was not rendered true in the context, (15). in the other baseline, the context again involved several individuals but the embedded clause was not rendered true in the context, (16). the experiment involved a latin square design so that each participant saw one condition from each item set for a total of 20 trials. (13) non-varying indefinite context—non-varying: [there were three homework problems to review. as a teaching assistant, ann was asked to review the homework problems. indeed, ann made sure all of the homework problems were error-free.] a teaching assistant made sure that every problem was error-free. (14) varying indefinite context—varying: [there were three homework problems to review. as teaching assistants, ann, bea and carol were asked to review the homework problems. ann made sure that the first problem was error-free. bea made sure that the second problem was error-free. carol made sure that the third problem was error-free.] a teaching assistant made sure that every problem was error-free. (15) non-varying indefinite context—non-varying baseline: [there were five homework problems to review. as a teaching assistant, ann was asked to review the homework problems. ann made sure that three of the five homework problems were error-free. however, she didn’t make sure that the fourth and fifth homework problems were error-free.] a teaching assistant made sure that every problem was error-free. (16) varying indefinite context—varying baseline: [there were five homework problems to review. as teaching assistants, ann, bea and carol were asked to review the homework problems. ann made sure that the first problem was error-free. bea made sure that the second problem was error-free. carol made sure that the third problem was error-free. none of them made sure that the fourth and fifth problems were error-free.] a teaching assistant made sure that every problem was error-free. for the experiment on cumulation with plural subjects, each item set also comprised of four conditions. the target sentence in each condition involved either a singular/plural subject, an embedding predicate and a tensed complement clause containing a universal quantifier, and the context always varied. in one condition (i.e., the control), the context involved a single individual that rendered the embedded clause true, (17). in another condition, the context involved two individuals, so that the embedded clause was only rendered true under a cumulative construal (by combining the contributions of each individual), (18). the other two conditions were unacceptable baselines. in one baseline, the context again involved a single individual but the target sentence involved a conproceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 282 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ joined subject so that there was a mismatch between the number of individuals in the context and the number of individuals in the target sentence, (19).2 in the other baseline, the context involved several individuals but the target sentence now involved a singular subject so that there was a mismatch between the number of individuals in the context and the number of individuals in the target sentence, (20). the experiment involved a latin square design so that each participant saw one condition from each item set for a total of 20 trials. at the end of both experiments, participants filled out the same post experiment questionnaire as in the first experiment on buildup cues. (17) singular subject, singular context—singular: [there were four homework problems to review. ann made sure that all the homework problems were error-free.] ann made sure that every problem was error-free. (18) plural subject, plural context—plural: [there were four homework problems to review. ann and bea, who don’t know each other, worked separately to make sure they were error-free. ann made sure that the first and second problems were error-free, but she didn’t look at the third and fourth problems. bea made sure that the third and fourth problems were error-free, but she didn’t look at the first and second problems.] ann and bea made sure that every problem was error-free. (19) plural subject, singular context—plural baseline: [there were four homework problems to review. ann made sure that all the homework problems were error-free.] ann and bea made sure that every problem was error-free. (20) singular subject, plural context—singular baseline: [there were four homework problems to review. ann and bea, who don’t know each other, worked separately to make sure they were error-free. ann made sure that the first and second problems were error-free, but she didn’t look at the third and fourth problems. bea made sure that the third and fourth problems were error-free, but she didn’t look at the first and second problems.] ann made sure that every problem was error-free. 4.2. results. the results for experiment 2 are provided in figure 2. the main observation is that, in the plot on the left, the plural condition is rated better with cumulating predicates than with non-cumulating predicates. similarly, in the plot on the right, the varying condition is also rated better with cumulating predicates than with non-cumulating predicates. these results suggest that cumulation with plural subjects and variation readings with singular indefinites show the same predicate sensitivity. the results were analyzed using linear mixed effects models (using the lme4 package in r) with predicate type and context as fixed effects—including interactions—and random intercepts and slopes by participant (including interactions) and random intercepts and slopes for context by item (since predicate type did not vary by item).3 the statistical models indicate 2it is worth pointing out that while this was meant to serve as an unacceptable baseline, some participants rated these sentences fairly high. we speculate that the reason is that, even though a single individual rendered the embedded clause true in the context, participants were able to interpret the target sentence collectively, as if the conjoined subject in the target sentence formed a group of some kind. while this was not expected, we believe this is an interesting result in its own right and worth exploring more in future research. 3model specification to analyze results: (i) lmer(response ∼ noncumulating.vs.cumulating*varying.vs.not + proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 283 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ singular plural plural baseline singular baseline singular plural plural baseline singular baseline 0 2 4 6 re sp on se acceptability ratings for cumulation non−varying varying varying baseline non−varying baseline non−varying varying varying baseline non−varying baseline 0 2 4 6 re sp on se acceptability ratings for (non)−variation readings predicate type cumulating non−cumulating figure 2: cumulation/variation readings with embedded universal quantifiers. left: cumulation with singular/plural subjects. right: variation readings with singular indefinites. that there was a significant interaction between predicate type and context for both cumulation with plural subjects (β = -0.916, p <0.001) and variation readings with singular indefinites (β = -0.649, p <0.05). we note that that the varying condition with non-cumulating predicates was not rated as low as the unacceptable baselines. this could be i) because the infelicity of the available reading and the actual context is more subtle than in our baseline condition, or ii) it could mean that a variation reading is harder with non-cumulating predicates but still possible (through cumulation or some other route). that being said, there is still a relative contrast between the varying conditions with cumulating and non-cumulating predicates. the positive results obtained in these experiments also provide support for the methodology that we used, namely, acceptability rating tasks. our results illustrate that, even for an experiment this size, this methodology is in fact sensitive enough to detect relative contrasts in acceptability for subtle judgments, like variation readings. in the above model, we divided the predicates into cumulating and non-cumulating ones based on intuitions concerning which predicates allow cumulation. to compare both kinds of predicates, we averaged over all non-cumulating predicates—even though there was variability in the degree to which each predicate licensed variation readings. this variability suggests that cumulation and variation readings may be gradient phenomena. however, the results from the experiment concerning cumulation validate this intuition by giving us a way to quantify the degree to which an embedding predicate allows cumulation (henceforth, ‘cumulation score’). this means we can test our research hypothesis without relying on untested intuitions by testing whether cumulation score predicts to what extent a variation reading is allowed (the prediction: a higher cumulation score leads to a better variation reading). to this end, we ran a second model, where the model specification was the same, but we modified how we coded the fixed effect of predicate type. using the results from the experiment on cumulation with plural subjects, we assigned each predicate a cumulation score by, first, looking at the ‘plural subject, plural context’ condition and calculating the average acceptability rating for each predicate in this condition. we then did the same for (1+varying.vs.not || item) + (1+noncumulating.vs.cumulating*varying.vs.not || participant) proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 284 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the ‘singular subject, singular context’ condition. we then took our cumulation score to be the difference between these two values. by doing this, we were able to model the degree to which a predicate allows cumulation (as a continuous predictor), instead of assuming it is a binary predictor in our model (cumulating vs. non-cumulating). the statistical model again indicates a significant interaction between predicate type and context (β = -0.763, p <0.05). furthermore, cumulation score does a fairly good job of dividing the predicates into those that we referred to as cumulating and those which we referred to as non-cumulating, by assigning the former a higher cumulation score. the division wasn’t perfect but, in general, reflected our initial intuitions.4 5. discussion. the results from the experiment on buildup cues provide evidence that variation readings are not easily available for both non-buildupicle predicates and (crucially) buildupicle predicates, even when external buildup cues were used. more specifically, according to the results from that experiment, a variation reading was not available for the crucial data point in (4) (compared to other non-buildupicle predicates). these null results provide evidence that, contra hoeks et al. 2022, the availability of variation readings is not mediated by the event structure of the matrix event, and furthermore that non-local qr may not be the mechanism underlying variation readings. in contrast, the results from the two experiments on cumulation with plural subjects and variation readings with singular indefinite subjects (from section 4.2) show a correlation between the degree to which a predicate allows for cumulation and the degree to which it licenses variation readings, which is as predicted by the cumlulating approach, but not predicted by the buildup approach. these results provide evidence for the empirical generalization in (21): (21) cumulation-variation correspondence: an embedding predicate licenses variation readings (apparent wide scope of a universal) whenever it licenses cumulation with plural subjects. the cumulating approach dispenses with the need for imposing a buildup constraint on non-local qr (and also the need for non-local qr out of tensed clause complements in general). that being said, the cumulating approach still captures the intuition that the truth conditions of examples like (2) involve adding up individual cases of making sure toward the overall reading. in conclusion, the results from these two experiments suggest variation readings are licensed by an embedding predicate’s inferential/cumulating properties, and not simply the event structure of the matrix event. as a result, we can maintain that tensed clauses are scope islands for universal quantifiers after all and that apparent wide scope of the universal is derived indirectly via cumulation. 6. variation readings with negative quantifiers. in section 4, we presented an argument in favour of the cumulating approach, namely, the correlation between the degree to which a predicate allows for cumulation and the degree to which it licenses variation readings. we now turn to a second argument in favour of the cumulating approach, put forth in palucci 2024. consider (22), where the embedded universal quantifier is replaced by a negative quantifier (i.e., no problem). (22) [there were three homework problems to review. as teaching assistants, ann, bea and carol were asked to review the homework problems. ann made sure that the first problem 4there were some exceptions. the predicate note down, which we classified as cumulating based on intuitions, received a rather low cumulation score. furthermore, some predicates which we classified as non-cumulating patterned more like cumulating predicates (realize, fear and confess) in terms of receiving higher cumulation scores. proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 285 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ didn’t contain errors. bea made sure that the second problem didn’t contain errors. carol made sure that the third problem didn’t contain errors.] a teaching assistant made sure that no problem contained errors. the observation from palucci 2024 is that (22) also licenses a variation reading where the teaching assistant can vary by problem. the challenge is that qr only delivers the right truth conditions when the embedded quantifier is a universal quantifier. even if the negative indefinite were to undergo qr to a position above the subject indefinite, as illustrated by the lf in (23-a), the resulting truth conditions correspond to an unattested reading. these truth conditions are provided in (23-b), where we can observe that the resulting truth conditions don’t correspond to a variation reading. (23) a. [no problem] λ1 [a teaching assistant made sure that tp[t1 contained errors]] b. ¬∃y [problem(y) ∧ ∃x[ ta(x) ∧ make-sure(x, contained-errors(y)) ] ] ‘there’s no problem y s.t. there’s a teaching assistant that made sure y contained errors.’ furthermore, palucci 2024 reports, not only are variation readings licensed with negative quantifiers, but both kinds of sentences (those with embedded universal quantifiers and those with embedded negative quantifiers) pattern in a similar way and show the same predicate sensitivity. if this claim is true, it would be strong evidence in favour of the cumulating approach since a qr based approach can’t even deliver the right truth conditions with embedded negative quantifiers. to test this claim, we replicated the experiment in palucci 2024. once again, this was an improved experiment with better stimuli, better experimental design, more predicates and more participants. 6.1. experimental design. to test whether embedded negative quantifiers pattern like embedded universal quantifiers and exhibit the same predicate sensitivity, we conducted 2 sentence rating experiments. the first experiment (looking at cumulation with plural subjects) involved 38 participants. the second experiment (looking at variation readings) involved 32 participants. similar to the previous experiments, participants were recruited on prolific., with the same restrictions as in the first two experiments. the experimental design paralleled that of experiment 2 except all target sentences contained an embedded negative quantifier instead of a universal quantifier. 6.2. results. the results for experiment 3 are provided in figure 3. the takeaway is that, while negative quantifiers seem to pattern in a similar way to universal quantifiers regarding predicate sensitivity, it is difficult to draw any conclusions concerning the interaction between predicate type and context. this is because acceptability ratings with non-cumulating predicates seem to be lower across the board: both control items (i.e., singular condition for cumulation and non-varying condition for variation readings) had lower acceptability ratings with non-cumulating predicates than with cumulating ones. it could be that participants simply dis-preferred the non-cumulating predicates, compared to the cumulating ones. in fact, upon further analysis, statistical models indicated no significant interaction between predicate type and context for both cumulation with plural subjects (β = -0.409, p >0.05) and variation readings with singular indefinites (β = -0.254, p >0.05). however, the results of this experiment do provide evidence that variation readings are available with negative quantifiers, just that the predicate sensitivity may be different than with universal quantifiers. if so, this suggests that mechanisms other than qr, which cannot deliver the proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 286 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ variation reading, must be available—even if it is not clear whether cumulation is that mechanism.5 singular plural plural baseline singular baseline singular plural plural baseline singular baseline 0 2 4 6 re sp on se acceptability ratings for cumulation non−varying varying varying baseline non−varying baseline non−varying varying varying baseline non−varying baseline 0 2 4 6 re sp on se acceptability ratings for (non)−variation readings predicate type cumulating non−cumulating figure 3: cumulation/variation readings with embedded negative quantifiers. left: cumulation with singular/plural subjects. right: variation readings with singular indefinites. in sum, the experimental results with negative quantifiers don’t provide support for the cumulating approach but they do provide evidence that mechanisms other than qr are needed to account for variation readings, and as such, these results fit into the larger research program of identifying ‘pseudo-scope’ mechanisms (i.e., mechanisms which derive similar truth conditions as qr but don’t involve covert scope shifting/covert movement) (fox & sauerland 1996). 7. concluding remarks. in this paper, we focused on the relative scope between an embedded universal dp and a singular indefinite in the the matrix clause, following barker 2022, hoeks et al. 2022. however, when looking at scope-taking over expressions other than indefinites, it seems that wide scope from within a tensed clause is impossible. consider (24), taken from palucci 2024. (24) [the race can only have one winner.] #i consider it possible that every runner will win. in principle, there are two readings of (24). the first is a surface scope reading, where there is more than one winner and everyone wins. this reading is ruled out by the scenario though. the second is the inverse scope reading, where each runner has a chance at being the winner: for each runner x, i consider it possible that x wins. this reading is compatible with the scenario. if non-local qr is possible, the inverse scope reading should be available and the sentence should be felicitous— contrary to fact. this suggests that the inverse scope reading is not attested. again, we can make sense of this if, contrary to apparent counterexamples, tensed clauses may be scope islands for qr after all, and apparent counterexamples are due to other mechanisms, such as cumulation. 5in fact, we even have a second reason to doubt that variation readings with negative quantifiers involve the same underlying mechanism as variation readings with universal quantifiers. consider the contrast between (i-a) and (i-b). (i) a. a different student made sure that every speaker had a ride. b. #a different teaching assistant made sure that no problem contained errors. the relevant observation is that (i-a) licenses an internal reading of the adjective different, while (i-b) doesn’t. if the same underlying mechanism derived variation in both cases, prima facie, we would expect both sentences to license internal readings of different, contrary to fact. one possibility we consider for future research is that variation readings in the cases with embedded negative quantifiers arise due to the singular indefinite subject denoting an individual concept (a function from worlds/situations to individuals), instead of an existential quantifier. proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 287 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ references anderson, catherine. 2004. the structure and real-time comprehension of quantifier scope ambiguity: northwestern university dissertation. barker, chris. 2022. rethinking scope islands. linguistic inquiry 53(4). 633–661. https://doi.org/10.1162/ling a 00419. chomsky, noam. 1975. questions of form and interpretation. in scope of american linguistics, 159–196. berlin: de gruyter mouton. https://doi.org/10.1515/9783110857610-008. farkas, donka f & anastasia giannakidou. 1996. how clause-bounded is the scope of universals? in semantics and linguistic theory (salt) 6, 35–52. https://doi.org/10.3765/salt.v6i0.2764. fox, danny & uli sauerland. 1996. illusive scope of universal quantifiers. in north east linguistics society (nels) 26, 71–85. harada, masashi. 2022. locality effects in composition with plurals and conjunctions: mcgill university, montreal dissertation. hoeks, morwenna, deniz özyıldız, jonathan pesetsky & tom roberts. 2022. event plurality & quantifier scope across clause boundaries. in semantics and linguistic theory (salt) 32, 443–462. https://doi.org/10.3765/salt.v1i0.5393. may, robert. 1977. the grammar of quantification.: massachusetts institute of technology dissertation. palucci, jonathan. 2024. pseudo-scoping out of tensed clauses: the case of cumulation. in north east linguistics society (nels) 54, . syrett, kristen. 2015. experimental support for inverse scope readings of finite-clauseembedded antecedent-contained-deletion sentences. linguistic inquiry 46(3). 579–592. https://doi.org/10.1162/ling a 00194. wurmbrand, susanne. 2018. the cost of raising quantifiers. glossa: a journal of general linguistics 3(1). https://doi.org/10.5334/gjgl.329. proceedings of elm 3: 276-288, 2025 jonathan palucci: pseudo-scoping out of tensed clauses: cumulation vs. buildups. 288 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ less-comparatives must be less ambiguous than exactly-differentials, experimental data shows fabian schlotterbeck & polina berezovskaya* abstract. scope mobility of comparative operators has been claimed to surface in a narrow class of specific cases where intensional verbs are combined with lesscomparatives or exactly-differentials. though not uncontroversial and dependent on subtle judgments, this type of ambiguity influenced subsequent compositional semantic analyses of comparatives and was also used as a diagnostics for scope mobility of the comparative operator in cross-linguistic studies. we use judgment data from three acceptability rating experiments to empirically test the (un)availablity of this ambiguity in german and english. we discover an empirical difference between exactlydifferentials and less-comparatives which is unexpected under the standard approach to the semantics of comparatives. we discuss the theoretical implications of our findings and highlight recent proposals that can account for our data. keywords. scope ambiguity; degree semantics; comparatives; modals; acceptability judgments 1. introduction. it has been a little over twenty years since the seminal paper by heim (2000, following the standard degree theory by stechow 1984) broached the issue of degree operators and scope. according to heim (2000) and contrary to previous conclusions, e.g. from the comprehensive analysis of kennedy (1997), scope mobility of comparative operators surfaces in a narrow class of specific cases where intensional verbs are combined with less-comparatives or exactly-differentials, as in (1). according to this view, (1) has the two readings in (1-a/b). its less prominent, inverse scope reading in (1-b) imposes no upper limit on the paper’s length, but only a minimal requirement of 15 pp. (1) (this draft is 10 pages.) the paper is required to be exactly 5 pages longer than that. a. linear scope: ∀w ∈ acc : max{d : longw(p, d)} = 15pp ‘it is required of the paper that it is exactly 15pp long.’ b. inverse scope: max{d : ∀w ∈ acc : longw(p, d)} = 15pp ‘the minimum length required for the paper is exactly 15 pages.’ (where acc is the set of accessible worlds) though not uncontroversial and dependent on notoriously subtle judgments, this type of ambiguity influenced subsequent compositional semantic analyses (e.g., bhatt & pancheva 2004, breakstone et al. 2011, lassiter 2012, beck 2012a, a.m.o.) and also served as a diagnostics for scope mobility of the comparative operator in cross-linguistic research (e.g., beck et al. 2004, 2009). *this research was funded by the deutsche forschungsgemeinschaft (dfg, german research foundation) – project id 75650358. we would like to thank sigrid beck, doris penka and paula menéndez benito for valuable feedback on this project. we would also like to thank tamara stoyke and robin edds for their assistance in generating english materials and hening wang for support with the application for ethics approval. last but not least, we thank anonymous reviewers for comments on earlier versions of this paper. affiliation: university of tübingen; email: {fabian.schlotterbeck, polina.berezovskaya}@uni-tuebingen.de). proceedings of elm 3: 344-358, 2025 c©2025 fabian schlotterbeck and polina berezovskaya published by the lsa with permission of the author(s) under a cc by license. 344 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ in the present paper, we investigate the empirical status of this rather subtle ambiguity for german and english. in particular, we systematically collected judgments regarding the availability of the inverse scope reading in three acceptability rating experiments. this embeds our paper in the more general efforts of strengthening the database for theoretical linguistics (for discussion see, e.g., gibson & fedorenko 2013, gibson et al. 2012, sprouse & almeida 2013). 2. theoretical background. it is common in the semantic literature to analyze degree operators like the comparative, but also their class mates superlatives, equatives or measure phrases, etc., to be quantifiers over degrees (stechow 1984, heim 1985, 2000, beck 2011; a.m.o.). however, we find far less scope interaction than expected between these degree operators and other scope-bearing elements. this led alternative analyses, such as kennedy (1997), to take a nonquantificational approach, in which the comparative is not a quantifier and cannot undergo quantifier raising (qr). in response, heim (2000) explains the scarcity of scope interactions involving degree phrases (degps) with a number of independently motivated constraints (for further discussion on the constraints of quantifier movement of comparatives, cf. beck 2011; 1363). with stateva (2000), heim (2000) concludes that inverse scope can be detected when intensional verbs (like require, need or allow) scopally interact with non-monotone (e.g. exactly-differentials; as in (1)) or downward-monotone (e.g. less-comparatives; as in (2)) degps. (2) (context: the draft is 10 pages.) the paper is required to be less long than that. a. ∀w ∈ acc :max(λd. the paper is d-long in w) < 10pp ‘it is required of the paper that it is less than 10 pages long (and not longer).’ b. max(λd.∀w ∈ acc : the paper is d-long in w) < 10pp ‘the minimum length required for the paper is less than 10 pages.’ we illustrate the derivation of both readings in a qr-based approach in (3).1 under the inverse scope reading, the comparative operator takes scope above the modal verb. for the exactlydifferential in (1), this leads to a reading where no upper limit is imposed on the paper’s length but only a minimal requirement of 15 pp in our example. the same general point can be made regarding less-comparatives as in (3). from a theoretical point of view (via heim 2000), no difference is thus expected with respect to the investigated ambiguity when comparing less-comparatives and exactly-differentials in english. the same goes for german, as discussed in the next section. 2.1. importance of the ambiguity in cross-linguistic research. under the assumption that degps are generalized quantifiers over degrees, the illustrated ambiguity is used in beck et al. (2004, 2009) to argue for the existence of degree abstraction cross-linguistically. an additional diagnostic are negative island effects. german is on a par with english in displaying both, the purported ambiguity and negative island effects. it has therefore been taken to have the positive setting of the degree abstraction parameter (dap). the parameter asks whether any given language has binding of degree variables in the syntax. essentially, the question is whether logical forms with the following constellation exist: [ degp⟨⟨d, t⟩, t⟩ [ λd. [...td...]]]. 1in the logical forms, we use intensional meanings, i.e. propositions of type ⟨s, t⟩ where needed. modals like ‘require’are treated as quantifiers over possible worlds. these assumptions do not have an impact on our subject matter. 2 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 345 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (3) a. t ⟨⟨s, t⟩, t⟩ required ⟨s, t⟩ λti t degpj ⟨⟨d, t⟩, t⟩ op than that ⟨d, t⟩ λtj ⟨t⟩ the paper is ti to be tj longer (op ∈{exactly 5pp comp, less}) b. t degpj ⟨⟨d, t⟩, t⟩ op than that ⟨d, t⟩ λtj t ⟨⟨s, t⟩, t⟩ required ⟨s, t⟩ λti t the paper is ti to be tj longer it is worth mentioning that there is a substantial body of work on degree constructions in a range of individual languages spurred by this line of research. the suggested diagnostics have been used for cross-linguistic (field)work and l1 acquisition in a range of languages (see e.g. bochnak 2015 on washo; bowler 2016 on warlpiri; hohaus et al. 2014 for l1 acquisition of german and english; howell (2012) on yorùbá; berezovskaya (2014) on l1 acquisition in russian; and kapitonov (2019) on kunbarlang). among the 17 languages from different language families investigated by beck et al. (2009), motu, japanese, chinese, mooré, samoan and yorùbá were determined to have the negative setting of the dap-parameter. german patterned with english, bulgarian and hindi-urdu in terms of the positive parameter setting of the dap. here, no differences are expected between english and german with respect to our protagonist, the modal-comparative ambiguity. combining this with prominent theoretical proposals discussed above, we can thus derive the prediction for the current experiments that the inverse scope reading should be available in german and english for both exactly-differentials and less-comparatives. a limitation of the cross-linguistic research by beck et al. (2009) for our current purpose is, however, that their questionnaires included few items and no inferential statistics were reported regarding the comparison between exactly-differentials and less-comparatives or a baseline condition. thus, we cannot draw strong conclusions about the existence of the inverse scope reading from this research. interestingly, recent work by philipp & zimmermann (2020, 2023) points to a gradual, non-categorical difference between english and german with regard to scope ambiguity in nominal quantification, with higher acceptance of inverse scope in english. from this perspective, it is an interesting question whether there is a difference in the scope potential of the comparative operator between english and german as well. 3. experiments on german. in order to test for the availability of inverse scope readings in sentences like (1) and (2) in german, we conducted two web-based questionnaire studies. 3 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 346 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3.1. experiment 1. the aim of exp. 1 was to compare judgments indicating linear or inverse scope readings of sentences like (1) and (2) to a baseline obtained from judgments for simple unambiguous control conditions. 3.1.1. methods. materials, design, procedure & participants: twelve german items were constructed as exemplified in (4) and (5). all items start with a sentence in which gradable adjectives (e.g. lang, ‘long’) are degree-modified by exactly-differentials (e.g. genau 10 seiten länger als, ‘exactly 10 pages longer than ...’) or less-comparatives (e.g. weniger lang als, ‘less long than ...’).2 gradable properties of referents (e.g. ‘the draft’ and ‘the paper’) that are introduced by definite descriptions are compared. in half of the conditions (e.g. (4-a/c)), comparatives are combined with the modal verb müssen (‘must’). by hypothesis, the presence of ‘must’ leads to the purported ambiguity. sentences without modals (e.g. (4-b/d)) were used as unambiguous controls against which responses to the modal conditions can be compared (see predictions below). the controls were kept similar to the target conditions regarding aspects like syntactic structure, lexical material, length, plausibility, etc. at the same time, we kept them as simple as possible in order to maximize our chances of detecting the ambiguity. the rationale behind this was that comprehension difficulty and ambiguity may have indistinguishable effects on judgments. in particular, ambiguous as well as difficult sentences may lead to less clear-cut judgments than sentences that are easier to understand or unambiguous, respectively. if we found no indication of the ambiguity even when comparing the relatively complex test sentences to very simple controls (i.e. if response categories are differentiated to the same degree in both cases), we would thus have rather strong evidence against the purported ambiguity in the tested type of construction. to keep the exactly-controls as easy to understand as possible, we removed not only the modal, but also the comparative morphology from them. the reason was that previous studies found increased processing difficulty of comparative vs. bare forms of the adjective (e.g. agmon et al. 2019). this was not possible for the less-controls because they are inherently comparative, i.e. they need to have a standard of comparison. sentences in all conditions are followed by a short post-context sentence. (4) target sentences and post contexts a. das the papier paper muss must genau exactly 10 10 seiten pages länger longer sein be als than der the entwurf. draft. so so lautet sounds die the vorgabe guideline der of.the zeitschrift. journal ‘the paper is required to be 10 pages longer than the draft. that’s what the journal’s guideline says.’ b. das the papier paper ist is genau exactly 10 10 seiten pages lang. long. das that haben have die the autoren authors gesagt. said ‘the paper is exactly 10 pages long. that’s what the authors said.’ c. der the entwurf draft muss must weniger less lang long sein be als than das the papier. paper. so so lautet sounds die the vorgabe guideline der of.the zeitschrift. journal ‘the draft is required to be less long than the paper. that’s what the journal’s guideline says.’ d. der the entwurf draft ist is weniger less lang long als than das the papier. paper. das that haben have die the autoren authors gesagt. said ‘the draft is less long than the paper. that’s what the authors said.’ 2materials, experimental data and analysis scripts of the experiments can be found in the following repository in the open science framework: https://osf.io/528cr/?view only=a0c996854f75430585eeace550d3d01b 4 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 347 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ each sentence doublet is, furthermore, paired with yes-no comprehension questions, as illustrated in (5). there are two types of questions: “matching” questions probe for the preferred or (in case of the controls) only possible reading. “mismatching” questions ask about propositions that are incompatible with the linear scope readings and would thus receive a “no”-response if this was, in fact, the only available reading. a “yes”-response to a mismatching question is indicative of an available inverse scope reading in the modal conditions or an error, e.g. in comprehension or response execution, in the controls. the pairing of target sentences and comprehension questions is indicated by the labels in (4) and (5); e.g. (4-a) is paired with the matching question in (5-a-m) and mismatching question in (5-a-mm). (5) comprehension questions a-m soll should das the papier paper 10 10 seiten pages länger longer sein be als than der the entwurf? draft ‘should the paper be 10 pages longer than the draft?’ a-mm darf may das the papier paper auch also 15 15 seiten pages länger longer sein be als than der the entwurf? draft ‘is the paper also allowed to be 15 pages longer than the draft?’ b-m ist is das the papier paper 10 10 seiten pages lang? long ‘is the paper 10 pages long?’ b-mm ist is das the papier paper 14 14 seiten pages lang? long ‘is the paper 14 pages long?’ c-m soll should der the entwurf draft kürzer shorter sein be als than das the papier? paper ‘should the draft be shorter than the paper?’ c-mm darf may der the entwurf draft auch also länger longer sein be als than das the papier? paper ‘is the draft also allowed to be longer than the paper?’ d-m ist is der the entwurf draft kürzer shorter als than das the papier? paper ‘is the draft shorter than the paper?’ d-mm ist is der the entwurf draft länger longer als than das the papier? paper ‘is the draft longer than the paper?’ altogether, we thus manipulated the factors modifier (exactly vs. less), modal (absent vs. present) and question (match vs. mismatch), yielding eight conditions in a 2 × 2 × 2 design. the complete set of experimental items comprised 96 pairs of assertions and questions distributed over eight lists using a latin square (together with 48 fillers). target and post-context sentences were presented together on a first display.3 the yes-no-question was presented on a subsequent 3sentences were presented in a self-paced manner and phrase by phrase using the moving window technique. the reason was that, with regard to the filler items, which were part of another unrelated experiment, we were interested in reading times, which are, however, irrelevant and uninformative here. 5 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 348 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ display in each trial. participants indicated their response by pressing one of two response keys. the response key assignment was counterbalanced between participants. before exclusion, 87 participants were recruited via prolific.co (mean age: 23 years; sd: 6; 66 female; 19 male). statistical analysis: a consequence of our design decision to keep controls simple was that we observed quasi-complete separation with the lowest error rates in control conditions being 1.1% and 2.3% (corresponding to an absolute frequency of one and two) errors in exps. 1 &. 2, respectively. because quasi-complete separation can make parameter estimates and resulting p-values unreliable, we analyzed the data using bayesian logistic regression (although we did not observe the usual symptoms of large estimates or standard errors). in particular, we computed bayesian logit mixed-effects models (implemented in the blme package) with weakly informative priors. we assumed normally distributed priors with a mean of 0 and standard deviation of 5 for the fixed effects, as suggested by clark et al. (2023). this analysis revealed exactly the same qualitative effects as our original logit mixed-effects models, which we provide in the osf repository. 3.1.2. predictions. according to heim (2000), the presence of the modal verb müssen (‘must’) makes the inverse scope reading available, although it may still be dispreferred. based on this hypothesis – and assuming the control conditions work as intended – a larger number of incongruent answers was, therefore, expected for the modal as compared to the control conditions, irrespective of the degree modifier (less vs. exactly). this expectation can be illustrated with a simple example: a pattern of results that would meet the expectation could consist of 10% incongruent answers in the controls (say 90% “yes” in matching and 10% “yes” in mismatching questions) because of errors vs. 30% incongruent answers in the potentially ambiguous conditions (say 80% “yes” in matching and 40% “yes” in mismatching questions) because of the availability of the inverse reading. more generally, we predicted a modal × question interaction in the relative frequency of “yes” responses, where the frequencies of responses in the two categories diverge less clearly between the two question types for controls than for the potentially ambiguous conditions. by contrast, proposals that assume no scopal flexibility of the comparative operator would not predict any difference between the pattern of responses to conditions with and without modals. a larger number of incongruent answers in the modal conditions may, however, still be explained as errors due to increased semantic and syntactic complexity. the most informative result here would be one without any difference between the modal and control conditions, as this would be hard to account for under the assumption of scopal flexibility of the comparative operator. 3.1.3. results. after applying predefined exclusion criteria, data from 62 participants were passed on to the statistical analysis. participants were excluded if they (i) took more than two standard deviations longer than the mean duration of an experimental session, (ii) had individual rts of more than 10 s, (iii) were below chance in answering comprehension questions in filler trials, (iv) indicated german was not their native language, or (v) missed more than one break (i.e. pressed the same button as in experimental trials instead of a designated button to end the break). descriptive results are shown in fig. 1. control conditions with the modifier exactly led to fewer errors on average (1%) than controls with less-comparatives (11.3%). in the modal conditions, questions matching the preferred linear scope interpretation were overwhelmingly answered with “yes” (94.1%) and mismatching questions with “no” (83.3%). 6 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 349 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ exactly less match mismatch match mismatch 0.00 0.25 0.50 0.75 1.00 question re l. fr eq . " ye s" ( an d ap pr ox . 9 5% c is ) modal absent present figure 1: relative frequency of “yes” responses across conditions in exp. 1 and approximate confidence intervals (cis). confidence intervals were computed for binomial probabilities but ignoring the repeated-measures design. the statistical analysis revealed a significant three-way interaction (z = −3.05, p = .002). to resolve the three-way interaction, we conducted separate analyses of the two modifiers. in the analysis of the exactly-conditions, we found the predicted modal × question interaction (z = 3.36, p = .001). this effect was due to the fact that there was less variation in judgments for conditions with modal absent vs. present. pairwise comparisons showed a significant effect of the modal in mismatching (z = 5.19, p < .001) but not in matching (z = −1.44, p = .15) conditions. this contrasts with the less-comparatives where no modal×question interaction was found (z = −0.32, p = .747). in addition, the expected main effect of question, i.e. more positive responses to matching vs. mismatching questions, was found for both modifiers (z = −7.84 and z = −9.26, for exactly and less, resp.). however, the effect of modal was not significant in either subset of the data. 3.1.4. discussion. indication of the purported ambiguity was limited to exactly-differentials. this difference between the two modifier types was surprising to us, as they are usually not distinguished in the theoretical literature with regard to their scope taking potential (but see the general discussion below for possible explanations of the difference). however, even in the exactly-differentials, the indication of the purported ambiguity that we found may, in fact, be an artefact driven by the simplicity of the exactly-controls rather than ambiguity in the corresponding modal conditions: the exactly-controls were the only conditions without comparatives. they were therefore relatively easy to judge and led to almost flawless performance. as a matter of fact, a post-hoc analysis supports this interpretation. in particular, we found a modifier type×question-interaction in controls (z = 3.062, p = .002), but not in the conditions with modals (z = −0.449, p = .653). these results stem from the fact that in the control conditions, response options were differentiated more clearly for the exactly-differentials (in the positive form 7 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 350 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ of the adjective) as compared to the less-comparatives (in the comparative form of the adjective), whereas no such interaction effect was found in the target conditions (all in comparative form). thus, we conclude at this point that the data of exp. 1 provide little evidence for the purported ambiguity, at least in german. 3.2. experiment 2. exp. 2 was a quasi-replication intended to test whether the results of exp. 1 reflect differences between the two types of comparatives or were due to specific characteristics of the controls, i.e. potential differences in comprehension difficulty of the comparative vs. the positive form of the adjectives. to this end, exactly-controls, e.g. (4-b), were put into comparative form, e.g. (5-e) – like all the other conditions. (5) modified target sentences and post context in exactly-control of exp. 2 e. das the papier paper ist is genau exactly 10 10 seiten pages länger longer als than der the entwurf. draft. das that haben have die the autoren authors gesagt. said ‘the paper is exactly 10 pages longer than the draft. that’s what the authors said.’ (6) modified comprehension questions in exactly-control of exp. 2 e-m. ist is das the papier paper 10 10 seiten pages länger longer als than der the entwurf? draft ‘is the paper 10 pages longer than the draft?’ e-mm.ist is das the papier paper 14 14 seiten pages länger longer als than der the entwurf? draft ‘is the paper 14 pages longer than the draft?’ 3.2.1. methods. except for the modified exactly-controls, the design, materials, procedure and statistical analysis of exp. 2 were identical to exp. 1. the predictions were also the same as in exp. 1. we used the same type of statistical analysis to test for those predictions. in total, 87 new participants were recruited via prolific.co (mean age: 33.17 years; sd: 11.2; 46 female; 41 male). 3.2.2. results. after applying the same exclusion criteria as in exp. 1, data from 61 out of 87 participants were passed on to the statistical analysis. mean judgments are depicted in figure 2. they are by and large comparable to the results from exp. 1, except for a bit more errors in the mismatching questions in the exactly-controls (exp. 1: 1%; exp. 2: 5.6%). there were, however, again fewer errors on average in the exactly-controls than in controls with less-comparatives (4% vs. 13.7%). in the modal conditions, matching questions were, as in exp. 1, overwhelmingly answered with “yes” (88%) and mismatching questions with “no” (92.9%). the statistical analysis revealed a significant three-way interaction (z = −2, p = .046). again, we resolved this interaction by using separate analyses of the two modifiers. in the analysis of the exactly-conditions, we found a marginal modal × question interaction (z = 1.86, p = .063), reflecting a larger difference between matching vs. mismatching questions in the modal than in the control conditions. pairwise comparisons showed a marginal effect of the modal within the two conditions with matching questions (z = −1.87, p = .061), with less variation in judgments for conditions with modal absent vs. present. no such effect was found in mismatching exactlyconditions (z = .68). as in exp. 1, this contrasts with the less-comparatives where no interaction 8 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 351 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ was found (z = −0.9, p = .38). furthermore, the expected main effect of question, i.e. more positive responses to matching vs. mismatching questions, was found for both modifier types (z = −8.15 and z = −9.8, for exactly and less, resp.). finally, there was an effect of modal that was only significant in less-comparatives (z = −2.97, p = .003). the latter effect was due to the fact that there were fewer “yes”-responses for modal than control conditions. this numerical pattern was also observed in exp. 1, but there it did not lead to a significant effect. exactly less match mismatch match mismatch 0.00 0.25 0.50 0.75 1.00 question re l. fr eq . " ye s" ( an d ap pr ox . 9 5% c is ) modal absent present figure 2: relative frequency of “yes” responses across conditions in exp. 2 and cis computed for binomial probabilities ignoring repeated-measures. 3.3. discussion. although the comparative exactly-controls in exp. 2 did in fact lead to a few more errors as compared to exp. 1, as we expected, the general pattern of results was the same in both experiments. indication of the ambiguity was again limited to the exactly-conditions. in these conditions, it hinges on a marginal difference in the generally rather high proportion of “yes”responses to the matching questions in exactly-controls vs. exactly-targets. a difference between the two experiments was a significant effect of modal in exp. 2 that was absent from exp. 1. given, its inconsistency, we cannot be confident that it is a real effect. moreover, this was an unexpected effect and we therefore refrain from speculating about potential reasons. 4. experiment 3 on english: english materials with contextual embedding. the first two experiments yielded consistent and highly comparable results. nevertheless, there are still questions regarding their generality. exp. 3 addressed four of these questions. the first question concerns the tested language. exps. 1 & 2 used german materials, although the original considerations of kennedy (1997) and heim (2000) were focused on english. while we have no specific reason to believe that there are relevant differences in the semantics of comparative operators between the two languages (see section 2 for a brief discussion), there is empirical evidence that scope ambiguities can be more pronounced in english as compared to german (e.g., philipp & zimmermann 9 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 352 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 2020, 2023). in exp. 3, we therefore tested english translations of the items from the previous two experiments. secondly, we embedded the translated items into contexts that were constructed in such a way as to make the inverse reading salient. the intention was to boost the potentially dispreferred inverse reading as much as possible. thirdly, we used demonstrative that in the thanclause to refer to the reference degree that was introduced in the context, as was also the case in the original examples discussed by heim (2000). if the than-clause contains a definite description, a de-dicto reading is possible (e.g. ...than the paper is required to be; cf stechow 1984). while we do not see how this could have weakend the indication of ambiguity in the observed judgments of exps. 1 & 2, it may still cause differences between the materials tested in these experiments and the examples discussed by heim (2000). finally, we removed the auch (‘also’) from the mismatching questions in the modal conditions because it may have triggered a presupposition that the preferred linear scope reading is true. 4.1. methods. the items from exps. 1 & 2 were translated into english and embedded into contexts that were intended to make the inverse reading salient. demonstrative that was included in the than-clause and auch (‘also’) was removed from the mismatching questions. the context and target sentences from an example item are shown in (7). the than-clause referred to a degree that was introduced in the context. the final two sentences of the context varied depending on the condition of the target sentence. (7) alex’s company has been assembling motorbikes, scooters and mopeds for well-known manufacturers for years. after a major refurbishment of the main production facility, the machinery on the assembly line needs to be recalibrated. in order to ensure that the minimum height standards for different vehicle types are met, alex and his engineers are consulting various designs provided by the manufacturers. ... a. ...the specifications from one particular manufacturer are especially puzzling. for instance, one of the mopeds is 90 cm high. the motorbike is required to be exactly 50 cm higher than that. that’s what the manufacturer specified. b. ...they are also trying to incorporate their customers’ feedback. one of the customers indicated that their moped is 90 cm high. the motorbike is exactly 50 cm higher than that. that’s what the motorcyclist said. c. ...the specifications from one particular manufacturer are especially puzzling. for instance, one of the motorbikes is 140 cm high. the moped is required to be less high than that. that’s what the manufacturer specified. d. ...they are also trying to incorporate their customers’ feedback. one of the customers indicated that their motorbike is 140 cm high. the moped is less high than that. that’s what the motorcyclist said. participants read the context on a first display and the target sentence on a second one. experimental items were distributed across eight lists as in exps. 1 & 2 and 18 filler items were added to each list. the exclusion criteria were identical to exps. 1 & 2 except for the fact that no rt-based criterion was used and participants were excluded if the ‘accuracy’ in their ratings on designated filler trials was more than 2 standard deviations away from the mean. apart from these differences, the materials and procedure were identical. we recruited 202 participants over prolific.co (mean age: 45.6 years; sd: 14.8; 121 female; 79 male). 4.1.1. results. after exclusion, data from 199 participants were passed on to the statistical analysis. mean judgments are shown in fig. 3. as compared to exps. 1 & 2, we observed slightly more variation in responses across the board in the current experiment (mean judgments ranged 10 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 353 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ between 8.6% and 91.4%). nevertheless, we found a similar pattern of results overall, as evidenced by a significant three way interaction (z = 2.02, p = .044) and also similar effects within exactlydifferentials and less-comparatives. in the less-comparatives, we found main effects of question type (z = −18.23, p < .001) and modal (z = −3.57, p < .001) again but no interaction between them (z = −.18, p = .86). in the exactly-differentials, we found a significant interaction between question type and modal (z = 2.86, p = .004). this was due to more “yes”-responses in the mismatching condition in the presence vs. absence of modals (z = 4.08, p < .001). in the matching question type, there was again no difference due to modals (z = −.04, p = .97). exactly less match mismatch match mismatch 0.00 0.25 0.50 0.75 question re l. fr eq . " ye s" ( an d ap pr ox . 9 5% c is ) modal absent present figure 3: relative frequency of “yes” responses across conditions in exp. 3 and cis computed for binomial probabilities ignoring repeated-measures. 4.1.2. discussion. contextual embedding led to a bit more variance in judgments as compared to exps. 1 & 2. this is likely due to contextual embedding and increased memory load resulting from the presentation on two consecutive displays. apart from this difference, the qualitative pattern of results was highly comparable between all three experiments. thus, the results from the previous two experiments seem neither specific to german comparatives nor attributable to missing contextual support. moreover, our findings cannot be attributed to de dicto readings of the definite expressions in the than-clauses in exps. 1 & 2 or to the presence of the presuppositional item auch (‘also’) in the mismatching questions. this is because both features were absent in the current experiment. next, we turn to the theoretical implications we draw from the three experiments. 5. general discussion. although the theory of natural language quantifiers comes in different flavours (e.g., barwise & cooper 1981, heim & kratzer 1998, barker 2002, sternefeld 2020), it has been extremely successful in the sense that its basic assumptions explain a large number of different linguistic phenomena across languages. the approach of heim (2000; drawing on stechow 1984, heim 1985; a.m.o.) that was tested in the current study can be viewed as an attempt to increase 11 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 354 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the coverage and test the explanatory power of quantifier theory by applying it to comparative expressions as well (see also lassiter 2012; for a similar attempt), at least in languages with degree abstraction (cf. beck 2009). however, despite the granted elegance and predictive power of this general approach, our experimental data indicate that with regard to comparatives, it has to be critically evaluated when confronted with systematically collected speaker judgments. in both experiments on german and the experiment on english, indication of the purported ambiguity was limited to exactly-differentials, and even in these cases, evidence for the minimal requirement reading is not particularly strong. we were able to replicate the results for english even after changing the dependent measure and fixing potential problems such as the possibility of de dicto readings. the question is what sets exactly-differentials and less-comparatives apart with respect to scope taking. one possibility could be their monotonicity properties (non-monotone vs. downward entailing). potential effects of monotonicity were discussed at length by heim (2000). she concluded that inverse scope can be detected with both non-monotone and downward entailing quantifiers. however, it is not clear how an explanation of our current findings in terms of monotonicity might look like. another difference between the two modifiers is that decompositional analyses have been suggested for less than (e.g., heim 2006) but we are not aware of similar proposals for exactly-differentials (beyond the obvious; cf. below). decompositional analyses have in fact been related to ambiguities in less-comparatives containing modals (e.g., heim 2006, rullmann & beck 1996). comparable to our results, the prevalence of such ambiguities was, however, challenged in an empirical investigation by beck (2012b) using german materials. furthermore, decompositional analyses would likely predict more rather than less potential for scope interaction in less-comparatives, as compared to other, non-decompositional approaches. one possible explanation of our data could be derived from oda (2008; chap. 2), who accounts for the ambiguity in (1) in terms of scope mobility of the differential measure phrase exactly 5 pp rather than the comparative operator. she considers the possibility that exactly-differentials allow for scope interaction, because the exactly-measure phrase contained in them is a generalized quantifier over degrees (type ⟨⟨d, t⟩, t⟩) and can move. she derives the relevant readings from this assumption. by contrast, less-comparatives are not ambiguous in the same way. in the domain of equatives, penka (2024) in a response to hohaus & zimmermann (2021) entertains a similar possibility: namely that the heim-ambiguity is due to the scope of genau (‘exactly’) rather than the scope of the equative operator. another promising avenue for explanation could potentially be derived from the following difference: exactly is assumed to trigger obligatory implicatures, for instance by insertion of a covert exhaustification operator (landman 1998, gajewski 2008). another related account is that of beck (2012a) who derives scope ambiguity of less-comparatives using an alternative-semantics along the lines of rooth (1992), but also uses the proposal of oda (2008) to explain ambiguity in exactly-differentials. with respect to scope ambiguity in exactlydifferentials, there is thus no substantial difference between the proposals of oda (2008) and beck (2012a). irrespective of the explanation that may turn out to be correct, our data indicate that the empirical basis for scope mobility of comparative operators is even more restricted than previously assumed. we suggest that scope mobility of comparative operators across languages should be subject of further empirical scrutiny. 12 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 355 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ references. agmon, galit, yonatan loewenstein & yosef grodzinsky. 2019. measuring the cognitive cost of downward monotonicity by controlling for negative polarity. glossa: a journal of general linguistics 4(1). 1–18. https:doi.org/10.5334/gjgl.770. barker, chris. 2002. continuations and the nature of quantification. natural language semantics 10. 211–241. https://doi.org/10.1023/a:1022183511876. barwise, john & robin cooper. 1981. generalized quantifiers and natural language. linguistics and philosophy 4(2). 159–219. https://doi.org/10.1007/bf00350139. beck, sigrid. 2009. positively comparative. snippets 20. 4–6. beck, sigrid. 2011. comparison constructions. in claudia maienborn, klaus von heusinger & paul h. portner (eds.), semantics: an international handbook of natural language meaning, vol. 2, 1341–1389. de gruyter. https://doi.org/10.1515/9783110255072.1341. beck, sigrid. 2012a. degp scope revisited. natural language semantics 20. 227–272. https://doi.org/10.1007/s11050-012-9081-6. beck, sigrid. 2012b. lucinda driving too fast again – the scalar properties of ambiguous thanclauses. journal of semantics 30(1). 1–63. https://doi.org/10.1093/jos/ffr011. beck, sigrid, svetlana krasikova, daniel fleischer, remus gergel, stefan hofstetter, christiane savelsberg, john vanderelst & elisabeth villalta. 2009. crosslinguistic variation in comparison constructions. linguistic variation yearbook 9. 1–66. https://doi.org/10.1075/livy.9.01bec. beck, sigrid, toshiko oda & koji sugisaki. 2004. parametric variation in the semantics of comparison: japanese versus english. journal of east asian linguistics 13(4). 289–344. https://doi.org/10.1007/s10831-004-1289-0. berezovskaya, polina. 2014. acquisition of russian degree constructions: a corpus-based study. in proceedings of formal approaches to slavic linguistics (fasl) 22, 1–22. bhatt, rajesh & roumyana pancheva. 2004. late merger of degree clauses. linguist inquiry 35(1). 1–45. https://doi.org/10.1162/002438904322793338. bochnak, ryan. 2015. the degree semantics parameter and cross-linguistic variation. semantics and pragmatics 8(6) 1–48. https://doi.org/10.3765/sp.8.6. bowler, margit. 2016. the status of degrees in warlpiri. in proceedings of the semantics of african, asian and austronesian languages (triple a) 2, 1–17. breakstone, micha y., alexandre cremers, danny fox & martin hackl. 2011. on the analysis of scope ambiguities in comparative constructions: converging evidence from real-time sentence processing and offline data. in proceedings of semantics and linguistic theory (salt) 21, 712–731. https://doi.org/10.3765/salt.v21i0.2609. clark, robert g., wade blanchard, francis k.c. hui, ran tian & haruka woods. 2023. dealing with complete separation and quasi-complete separation in logistic regression for linguistic data. research methods in applied linguistics 2(1). 1–11. https://doi.org/10.1016/j.rmal.2023.100044. gajewski, jon. 2008. more on quantifiers in comparative clauses. proceedings of semantics and linguistic theory (salt) 18. https://doi.org/10.3765/salt.v0i0.2494. gibson, edward & evelina fedorenko. 2013. the need for quantitative methods in syntax and semantics research. language and cognitive processes 28(1–2). 88–124. 13 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 356 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ https://doi.org/10.1080/01690965.2010.515080. gibson, edward, steven t. piantadosi & evelina fedorenko. 2012. quantitative methods in syntax/semantics research: a response to sprouse and almeida (2013). language and cognitive processes 28(3). 229–240. https://doi.org/10.1080/01690965.2012.704385. heim, irene. 1985. notes on comparatives and related matters. manuscript. austin: university of texas. heim, irene. 2000. degree operators and scope. in proceedings of semantics and linguistic theory (salt) 10, 40–64. https://doi.org/10.3765/salt.v10i0.3102. heim, irene. 2006. little. in proceedings of semantics and linguistic theory (salt) 16, 35–58. https://doi.org/10.3765/salt.v16i0.2941. heim, irene & angelika kratzer. 1998. semantics in generative grammar. malden: blackwell. hohaus, vera, sonja tiemann & sigrid beck. 2014. acquisition of comparison constructions. language acquisition 21(3). 215–249. https://doi.org/10.1080/10489223.2014.892914. hohaus, vera & malte zimmermann. 2021. comparisons of equality with german so. . . wie, and the relationship between degrees and properties. journal of semantics 38(1). 95–143. https://doi.org/10.1093/jos/ffaa011. howell, anna. 2012. comparatives and scope in yoruba: university of tübingen ma thesis. kapitonov, ivan. 2019. degrees and scales of kunbarlang. in proceedings of triple a5, 91–105. kennedy, christopher. 1997. projecting the adjective: the syntax and semantics of gradability and comparison: university of california at santa cruz dissertation. landman, fred. 1998. plurals and maximalization 237–271. dordrecht: springer netherlands. https://doi.org/10.1007/978-94-011-3969-410. lassiter, daniel. 2012. quantificational and modal interveners in degree constructions. in anca chereches (ed.), proceedings of semantics and linguistic theory (salt) 22, 565–583. https://doi.org/10.3765/salt.v22i0.2649. oda, toshiko. 2008. degree constructions in japanese. storrs: university of connecticut dissertation. penka, doris. 2024. german so ...wie-constructions as definite descriptions. a reply to hohaus & zimmermann (2021). manuscript. philipp, mareike & malte zimmermann. 2020. empirical investigations on quantifier scope ambiguities in german. in proceedings of sinn und bedeutung 24, 145–163. https://doi.org/10.18148/sub/2020.v24i2.914. philipp, mareike & malte zimmermann. 2023. an experimental comparison of the availability of inverse scope in english and german. linguistic inquiry 1–37. https://doi.org/10.1162/ling a 00493. rooth, mats. 1992. a theory of focus interpretation. natural language semantics 4(1). 75–116. https://doi.org/10.1007/bf02342617. rullmann, hotze & sigrid beck. 1996. degree questions, maximal informativeness, and exhaustivity. in paul dekker & martin stokhof (eds.), proceedings of the amsterdam colloquium 10, 73–92. sprouse, jon & diogo almeida. 2013. the empirical status of data in syntax: a reply to gibson and fedorenko. language and cognitive processes 28(3). 222–228. 14 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 357 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ https://doi.org/10.1080/01690965.2012.703782. stateva, penka. 2000. in defense of the movement theory of superlatives. in rebecca daly & anastasia riehl (eds.), proceeding of escol 1999, 219–226. clc publications. stechow, arnim von. 1984. comparing semantic theories of comparison. journal of semantics 3(1-2). 1–77. sternefeld, wolfgang. 2020. quantifiers, scope, and pseudo-scope. 1–44. john wiley and sons, ltd. https://doi.org/10.1002/9781118788516.sem124. 15 proceedings of elm 3: 344-358, 2025 fabian schlotterbeck and polina berezovskaya: less-comparatives must be less ambiguous than exactly-differentials, experimental data shows. 358 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ devoir, ou pouvoir, that is the question anouk dieuleveut & ira noveck* abstract. in languages like french and english, modals express either possibility (e.g., “you can”) or necessity (e.g., “you must”). previous acquisition research has shown that english-speaking children have particular difficulty with necessity modals: comprehension experiments show that they tend to accept must or have-to in possibility scenarios (noveck 2001, özturk & papafragou 2015, a.o.); production studies show that they use them less frequently than possibility modals, and when they do, their usage is not always adult-like (dieuleveut et al. 2022). but the cause of this “necessity gap” remains debated. one challenge is that past studies have focused primarily on english, where necessity modals are much rarer than possibility modals in parental speech, which could suggest that the delay is simply due to less exposure. in this study, we demonstrate through a corpus analysis of french young children’s modal use and their linguistic input, as well as experiments based on this data (following the methods of dieuleveut et al. 2022), that the delay cannot be attributed solely to limited exposure: despite more exposure, french-speaking children experience the same difficulties with necessity modals. furthermore, we show that these difficulties persist until children are five years old. keywords. modal acquisition; corpus study; human simulation paradigm; french/english comparison; necessity gap. 1. introduction. in languages like english and french, modals express either possibility or necessity. for instance, peux (‘can’), in (1a), means that it is possible for you to sleep—but you could just as well stay awake. dois (‘must’), in (1b), means that it is necessary for you to sleep—with no other option. in this paper, we investigate when and how children figure out the “force” of their modals: that pouvoir (‘can’) expresses possibility, whereas devoir (‘must’) expresses necessity. (1) a. tu peux dormir. ‘you can sleep’ possibility b. tu dois dormir. ‘you must sleep’ necessity the meaning difference between (1a) and (1b) appears obvious, but from the perspective of the child, figuring out modal force may not be so simple. indeed, the entailment relation between necessity and possibility creates a logical subset problem (berwick 1985, wexler & manzini 1987, a.o.). whenever a necessity modal statement like (1b) is true, the paired possibility statement (1a) is also true. if a learner mistakenly assumes that a modal means necessary, when it actually means possible, they will get evidence that their hypothesis is wrong, seeing it used in ‘possible but not necessary’ situations. however, if they assume a modal means possible when it actually means necessary, they won’t find direct evidence to falsify their hypothesis, as necessity always implies possibility. so what stops children from assuming that necessity modals like must mean possible? * we would like to thank morgan moyer, ailís cournane, valentine hacquard, annemarie van dooren, keny chatain, tom meadows, david mueller, dominique juffin, members of the grisp at the llf, members of the of the nyu language acquisition lab, audiences at the linguistic evidence conference in paris and audiences at bucld47. this project is supported by a fyssen foundation post-doctoral study grant and by the chaire d’excellence université de paris cité. authors: anouk dieuleveut, université de genève (anouk.dieuleveut@unige.ch) & ira noveck, laboratoire de linguistique formelle, umr 7110, université de paris, cnrs, france (ira.noveck@gmail.com). proceedings of elm 3: 130-143, 2025 c©2025 anouk dieuleveut and ira noveck published by the lsa with permission of the author(s) under a cc by license. 130 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ this problem, which arises for any pair of words that entertain entailment relations (like dog/animal, piantadosi 2011, xu & tenenbaum 2007, or some/every rasin & aravind 2021), is exacerbated by modal semantics. as modals are words used to talk about non-actual states of affairs, they lack clear reliable physical correlates—similarly to attitude verbs such as think or want (gleitman et al. (2005)’s “hard words”). moreover, in languages like french and english, modals can be used to express various types (or ‘flavors’) of modality: for instance, doit in (1b) can mean that you are required to sleep, but could also mean in a different context, that it’s likely. children have to acquire these two dimensions, force and flavor, in tandem. previous research has shown that english-speaking children struggle with necessity modals, drawing insights from both comprehension studies—which show that 4-year-olds tend to both over-accept possibility modals in necessity situations and necessity modals in possibility situations (noveck 2001, özturk & papafragou 2015, a.o.), and studies of their productions—which show that children start producing necessity modals later than possibility modals, use them less frequently, and when they do, do so in a non-adult-like way (dieuleveut et al. 2019, 2022). the origin of this “necessity gap” (dieuleveut 2021) is a matter of debate. children’s non-adult behavior has been attributed to a) conceptual difficulties reasoning with indeterminacy (cf the “premature closure” hypothesis from acredolo & horobin 1987; özturk & papafragou 2015, moscati 2017), b) semantic difficulties, with the meaning of necessity modals (children would not have figured out their underlying force) (dieuleveut 2021, cournane et al. submitted), or to c) pragmatic immaturity. but one issue limiting the conclusions we can draw is that these studies have focused primarily on english, where necessity modals are quite rare in the input as compared to possibility modals (dieuleveut et al. 2022). crucially, we do not know whether those difficulties are specific to english; the question of how children learn modal force has not been assessed in other languages. the goal of this study is to fill in this gap by comparing french to english, using the exact same methods in both languages. what makes this comparison particularly interesting is that we find in french parental speech the opposite pattern from english: necessity modals are more frequent. this allows us to directly test the effect of quantity of exposure on children’s mastery. we will show that despite hearing more necessity modals, french children face the same difficulties as their english counterparts: they produce necessity modals later, less frequently and tend to “overuse” them, i.e. to use them in situations where adults find possibility modals more appropriate. we conclude that the delay cannot be due only to low quantity of exposure. the rest of the paper is structured as follows. in section 2, we give a brief overview of modals’ semantics and review existing results from the acquisition literature, showing that english-speaking children struggle with necessity modals. we then turn to our study. first, in section 3, we report quantitative corpus data on french children’s modal productions and input, comparing our results to dieuleveut et al. (2022)’s for english.1 then, in sections 4 and 5, we report the two experiments run on the corpus data. the first one re-adapts dieuleveut et al. (2022)’s paradigm to french, allowing us to replicate their results. we extend the experiment to older children (4 to 5 y-o), where difficulties persist. the second experiment is a follow-up study carried out to exclude the possibility that our results stem from participants’ expectations that children use more possibility modals. we show they do not. in section 6, we reflect on the origin of children’s misuses, and identify questions that research should address next. additionally, we discuss our experimental paradigm’s advantages and limitations, and how it could be used for other cases of word learning. 1 we used the exact same methods as dieuleveut et al. (2022) for their corpus study on english. we present them on a par to ease comparison. preliminary results for french, with details about modals’ interaction with negation and flavor, can be found in dieuleveut (2023). proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 131 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 2. background. the meaning of modals is typically captured along two axes: force and flavor. the force corresponds to whether a modal expresses possibility (e.g., pouvoir ‘can’ in (1a)), or necessity (e.g. devoir ‘must’ in (1b)). in formal semantics, this difference is standardly captured by treating modals as quantifiers over (contextually determined) sets of possible worlds, paralleling quantifiers over sets of individuals (some/all). possibility modals express existential quantification (in some possible worlds, p is true). necessity modals express universal quantification (in all possible worlds, p is true). this analysis nicely captures the logical entailment between must and can: whenever (1b) is true, (1a) is also true. the flavor corresponds to the type of modality the modal conveys: possibility/necessity given what is known, as in (2a) (epistemic modality), or based on some rules, as in (2b) (deontic modality), or based on some goals (teleological modality) (kratzer 1981, 1991; see von fintel 2006, hacquard 2011, for overviews). typologies of modal flavors vary, but most authors agree on a major distinction between epistemic and non-epistemic, based on both syntactic and semantic criteria. here, we label non-epistemic as ‘root,’ following the terminology from hoffman (1966). (2) ira peut/doit dormir. ‘ira may/must sleep.’ a. epistemic: according to what we know, it is possible/certain/likely that… b. deontic (root): according to the rules, it is allowed/required that… c. teleological (root): according to his goals, it is possible/necessary to… in languages like french and english, modals (auxiliaries, semi-auxiliaries, as well as some adjectives and adverbs) lexically encode force, but vary in flavor.2 as highlighted in the introduction, modals thus raise a particularly complex learning problem. first, because the fact that necessity entails possibility creates a logical subset problem: given that every time they will encounter a necessity scenario, the possibility scenario is also true, what prevents learners from thinking that necessity modals just mean possibility? 3 with modals, this basic problem is complicated by the fact that modals lack clear reliable physical correlates, as discussed for attitude verbs (gleitman et al. (2005)’s “hard words”). plus, children must learn polysemous modal flavor— which might help or hinder force acquisition (see dieuleveut 2021, for discussion). behavioral experiments show that preschool-aged children have difficulty with modal force. 4-year-olds tend both to (i) over-accept possibility modals in necessity situations—e.g., accepting “there might be a bear in the box” in a situation where it is certain that there is one; and to (ii) accept necessity modals in possibility situations, where they are false—e.g., accepting “there has to be a bear in the box” when it is merely possible that there is one (noveck 2001, özturk & papafragou 2015, cournane et al. submitted, a.o.). the first result has received much attention in the context of children acquisition of scalar implicatures. indeed, the use of a possibility statement (like “there might be a bear in the box”) typically triggers a scalar implicature that the stronger statement doesn’t hold (as the speaker would otherwise have used it; see grice 1975). various studies have argued that children “difficulty” with implicatures depends on their ability to access alternatives (chierchia et al. 2001; barner et al. 2011, skordos & papafragou 2016, a.o.). they generally take for granted that children already know the meaning of scalar terms. but the second result, children’s acceptance of necessity modals in possibility situations, suggests it is not an innocuous step. 2 some other languages have modals with “variable force”, i.e., that can be used in situations where english speakers would use either a possibility or necessity modal (see e.g., deal 2011; see yanovich 2016 for a summary). 3 note that the statement ‘necessity entails possibility’ only holds keeping flavor constant. for instance, it is not the case that ‘it is certain that p’ entails ‘it is allowed’, or that ‘it is required that p’ entails ‘it is likely’. proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 132 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ studies of children’s modal productions often focus on the flavor dimension (see papafragou 1998, cournane 2020, for overviews). it has been shown in various languages that epistemic uses are “delayed”: children start producing root modals by age 2 (e.g., abilities, obligations), but produce epistemics only from around age 3 (the so-called epistemic gap, cournane 2015). corpus analyses focusing on force are rarer. one study on english (dieuleveut et al. 2019, 2022, based on the manchester corpus, theakston et al. 2001), using a combination of corpus analyses and experiments based on the corpus, reveals a “delay” in children’s mastery of necessity modals (necessity gap, dieuleveut 2021). their study shows that by the age of 2, children use possibility modals frequently, productively, and in an adult-like way, but they start using necessity modals later on, use them less often, and in a non-adult-like way: as when adults would expect possibility modals. to show this, they use a variant of the human simulation paradigm (gillette et al. 1999), where adult participants are asked to guess the force of modals uttered by children in dialogues extracted from the corpus. as they highlight in their discussion, their findings ground another hypothesis: that children are confused about the underlying force of their necessity modals—either they are merely uncertain or have encoded them with possibility force. as they point out, this explanation would also capture results from comprehension experiments with older children, which rarely ask whether children know the force of their modals. one crucial question is thus to determine whether children’s difficulty with necessity modals is specific to learners of english, or if it is more general. our study aims to address this question, comparing children’s developmental trends in french and in english using the exact same methods in both languages. one aspect that makes the comparison particularly relevant is the increased frequency of necessity modals in french children’s input, which allows us to directly assess the effect of quantity of exposure. 3. corpus study 3.1. methods. for french (fr), we use the lyon corpus (demuth & tremblay 2008) (5 childmother pairs; age range: 1;00-3;00; 3 females, 2 males) and the paris corpus (morgenstern & parisse, 2007) (6 child-mother pairs; 3f, 3m; age range: 0;7-6;03) (childes database, mac whinney 2000). for english (en), we use data from dieuleveut et al. (2022), based on the manchester corpus (theakston et al. 2001) (12 child-mother pairs; 6f, 6m; age range: 2;00-3;00). children were all recorded at home in unstructured play sessions with their parents. 3.2. coding. to stay close to dieuleveut et al. (2022) study, we focused on modal (semi)-auxiliaries and therefore excluded some other means to encode modality (adverbs like peut-être ‘maybe’, adjectives like possible, verbs like penser ‘think’ or vouloir ‘want’, modal verbal inflection (imperative, conditional, subjunctive). we excluded the semi-aux aller (‘go’), which (like will) conveys future. in french, pouvoir is the only modal expressing possibility; devoir, falloir, and avoir à all express necessity (chu 2008).4 all utterances containing modal (semi)-auxiliaries were extracted and coded for force (3), flavor (epistemic vs root) (4), complement (5) (french: adult: 5,231 utterances; 2-3-year-olds: 1,514; 3-5-year-olds: 1,404; excluding repetitions5: adult (2.2%): 5,114; 2-3-y-o: (8.9%): 1,379; 3-5-y-o: 1,296 (7.7%); english: adult: 20,755 utterances; child: 5,842; excluding repetitions and tag-questions: adult (9.1%): 18,853; child (17.8%): 4,800). 4 french modals are typically considered as semi-auxiliaries (hacquard 2010, borgonovo & cummins 2007). contrary to english auxiliaries, they inflect for tense, mood, aspect, and agree with their subject. note that falloir has a peculiar syntax: it only appears in impersonal constructions (with the expletive subject il) (*tu faux venir), and can take both infinitival and cp complements (il faut [venir]/il faut [que tu viennes]). 5 repetitions are cases where the speaker repeats a sentence uttered right before by herself or by another speaker. proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 133 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (3) functional modal lemmas by force: a. possibility: fr: pouvoir; en: can, could, might, may, able to b. necessity: fr: falloir, devoir, avoir à; en: must, should, need (to), have to, got to, supposed to, ought to (4) flavor: a. root: mother: ‘y a plein d'habits sales !’ (‘there are many dirty clothes!’) mother: ‘elle doit laver tout le linge.’ (‘she must do all the laundry’) b. epistemic: child: ‘je trouve pas la grosse.’ (‘i can’t find the big one’) mother: ‘elle doit être restée dans la voiture.’ (‘it must be in the car’) (5) complement: + np (excluded): fr: ‘il faut du pain’ en: ‘we need bread’ + inf: fr: ‘il faut venir’ (‘it is necessary to come’) + cp: fr: ‘il faut que tu viennes’ (‘it is necessary that you come’) 3.3. quantitative results. overall, utterances containing modals represent 3.8% of all adult utterances in french (5.8% in english), and 1.9% of child utterances (2.4% in english). as reported in other languages, epistemic uses are rare, both in adult and child speech (adults: fr: 5.9%; en: 8.8% of all modal uses; children: 2to 3-y-o: fr: 0.4%; en: 2.4%; 3to-5 y-o: fr: 1.8%; en: not assessed), except for devoir, which (like must) is more often used to convey epistemic than root modality (must: 64.2%; devoir: 60%) (see dieuleveut 2023). table 1 summarizes counts of french adult and child modal productions by force, with english as a comparison. in adult speech, we see a strong difference between french and english in the relative frequency of possibility and necessity modals, with french parents using necessity modals more often (62% of all their modal utterances, vs 28% in english). falloir is particularly frequent (53%). despite this difference in input, children in both languages produce necessity modals less often (fr: 38% of all modal utterances between 2 and 3; en: 21%), with only a slight increase for older children (42%). french (n = 7,379) english (n = 23,653) adults6 2-3-year-olds 3-5-year-olds adults 2-3-year-olds count (%mod utt) count (%mod utt) count (%mod utt) count (%mod utt) count (%mod utt) poss 2008 (38%) 850 (62%) 516 (58%) poss 13500 (72%) 3798 (79%) nece 3108 (62%) 529 (38%) 370 (42%) nece 5353 (28%) 1002 (21%) falloir 2659 (53%) 492 (36%) 298 (34%) devoir 403 (8%) 21 (2%) 66 (7%) avoir-à 46 (1%) 16 (1%) 6 (1%) all 5114 (100%) 1379 (100%) 886 (100%) all 18853 (100%) 4800 (100%) table 1: counts and percentages of modal uses by force and age group in french and english. frequency by lemma for english is available in dieuleveut et al. (2022). 3.4. age of first production. both french and english children start producing possibility modals quite early (first pouvoir/can: around 1 year 11 months), on average 4 months before their first necessity modals (first falloir/have to: around 2;03; devoir/must: 2;11; avoir-à: 5;06). 6 note that data for french adult talk corresponds to talk towards 2 to 3-yos, to make it comparable to the english study. however, there is not much variation in adult’s talk when they talk to older children (see appendix on https://osf.io/3cwqy/). we also do not see much variation between mothers, and no correlation between child and mother frequency of use–thought this might simply be due to the low sample size (11 children). note that since dieuleveut et al. (2022) includes children up to 3;3, the actual boundary between age groups is 3;03 and not 3;00. proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 134 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3.5. discussion. we find that both french and english children use necessity modals less frequently than possibility modals, and start using them later on, with a “gap” of almost 4 months. given the difference in their input, this is striking. however, sparse production does not imply lack of understanding. several factors might contribute to the prevalence of possibility modals in child early speech. first, differences in children and adults’ conversational status and goals: children might be less in a position to give orders or express certainty than adults, therefore less likely to use necessity modals. it could also be that young children avoid using necessity modals because they are harder to produce, or are cognitively costlier (there are few studies on processing of possibility and necessity modals even in adults). another factor that might play a role is the syntax of falloir, the most frequent french necessity modal, which could make it harder to acquire. beyond quantitative differences, dieuleveut et al. (2022) have shown, using a corpus-based experiment, that english children tend to use their necessity modals in a non-adult-like way: in situations where adults expect possibility modals (e.g., when adults expect “i can see it,” children say “i have to see it”). we adapted their method to french to see if we can generalize this result. 4. experiment 1. to get a finer-grained assessment of children’s uses, we used the method introduced by dieuleveut et al. (2019) (itself a variant of the human simulation paradigm, gillette et al. 1999).7 the goal is to determine whether children use their possibility and necessity modals in an adult-like way, i.e., in the same contexts as adults would, by asking adult participants to guess the force of a redacted modal uttered by children in dialogues extracted from the corpus. 4.1. methods. in the experiment, run online, adult participants read a series of mother-child dialogues randomly extracted from the corpus. their task is to guess the force of a blanked-out modal, by picking between two options, either a possibility (pouvoir) or a necessity modal (devoir; falloir). figure 1a illustrates a trial. we use participants’ accuracy in guessing the correct modal (uttered by children) as the measure of how ‘adult-like’ children’s uses were (indicating whether children used their modals in contexts where adults would). as a baseline, we use the same experiment on mothers’ modal utterances (figure 1b). procedure. all experiments were coded using penncontroller for ibex (zehr & schwarz, 2018) (https://www.pcibex.net/) and hosted on the llf ibexfarm server (https://ibex.llf-paris.fr/). overall, each participant had 40 dialogues to judge, presented in a randomized order: 20 controls using tense (past/future); 20 trials (10 possibility, 10 necessity, randomly selected out of a list of 20 dialogues randomly extracted from the corpus). a demo is available at https://ibex.llf-paris.fr/ibexexps/adieuleveut/elm3_demo/experiment.html (2 to 3 y-o, exp1d). conditions. we ran two versions varying the necessity modal (exp1d: devoir vs pouvoir; exp1f: falloir vs pouvoir; avoir à was too rare to be tested). we had three groups based on the speaker’s age: 2-to 3-year-olds, 4-to 5-year-olds, mothers (used as baseline). force was tested within subjects, age and lemma between subjects. we tested only ‘root’ modals because epistemic uses are too rare in children’s production, and we excluded negated utterances because of issues with scopal irregularities with negated modals (iatridou & zeiljstra 2013).8 for controls, participants had to pick between future and past (e.g. [a vu] vs. [va voir]). for these, we 7 the original goal of the human simulation paradigm is to compare different kinds of cues available to the child to figure out words’ meanings (gillette et al., 1999). here, we use with a different goal: as a way to evaluate children’s production. we use the hsp on adult’s production as a baseline, indicating that force can be guessed from conversational context. this method—using adults judgments to assess children usage—was developed in english (dieuleveut 2021, dieuleveut et al. 2022), and has also been used for definite description (the vs a) (ying et al., 2024). 8 dieuleveut et al. 2022 have two instantiations of root-positive (rootp1: can vs must; rootp2: can vs have to). they also test negated root modals (can’t vs don’t have to) and epistemic modals (might vs must). proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 135 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ensured the correct answer was not always guessable based on the target sentence alone, to force participants to read the entire dialogues. material. contexts were randomly selected from the corpus database for each modal lemma (pouvoir, devoir, falloir) (excluding negated cases, epistemic uses and repetitions) and manually checked. in cases where another modal was used in the dialogue (either by the child or the adult), we left it as such, regardless of its force. in cases where the same modal was used several times in the dialogue, we put several blanks. to replace falloir with pouvoir (in exp1f), we proceeded as follows: (i) for cases like ‘il faut [v+inf]’ (e.g., ‘il faut manger’ ‘(you) have to eat’), we used french impersonal subject on (participants had to pick between ‘il faut’ and ‘on peut’); (ii) for cases like ‘il faut [cp que tu manges]’, they had to pick between ‘tu peux manger’ and ‘il faut que tu manges’. when the expletive subject il was dropped, we put it back in the answers (e.g. ‘faut manger’ → ‘__ manger’: pick between ‘il faut’ and ‘on peut’). we used the same lists in exp1f and exp1d for possibility modals as a check, expecting no difference in accuracy. the final lists consisted of 20 dialogues for possibility and 20 dialogues for necessity for each three age groups (except for 2-3 y-o child devoir, where we could test only 17 contexts) (total: 237 test; 60 controls). they are available at https://osf.io/3cwqy/. enfant : ... t'en laisses un petit coup. maman : merci. enfant : voilà. maman : merci. enfant : arrête d'aller là avec le ptit chevaux enfant : vous arrêtez d'aller là. enfant : parce que c'est après. enfant : qu'on ________ aller après. doit peut autre adulte : oui. maman : oui autre adulte : une soucoupe. enfant : sont un peu vieilles. maman : oui sont un peu abîmées tordues. maman : ah celle-là elle marche bien. autre adulte : merci beaucoup. maman : tu ________ souffler dessus. dois peux child: ... you leave a little. / mum: thank you. / child: there you go. / mum: thank you. / child: stop going there with the horsie. / child: you stop going there. / child: because it's after. / child: that we ________ go after. other adult: yes. / mother: yes. other adult: a saucer. / child: they're a bit old. / mum: yes, they're a bit bent damaged. / mother: oh, this one works well. / other adult: thank you very much. / mother: you ________ blow on it. figure 1: example trials, experiment 1 (1d: pouvoir vs devoir). on the left: experiment on children’s production (2-3yo). on the right: experiment on mothers’ production (mot, baseline). 4.2. participants. 358 french participants were recruited on prolific (60 per condition, 2 failed to record data) (166 f, 186 m, 6 nb; mean age: 32.8yrs; age range: 18 to 74yrs). we removed 11 participants whose accuracy scores on controls were <75% (3.1%). we thus report results for 347 participants (exp1d: mot: 59; 2-3yo: 56; 4-5yo: 59; exp1f: mot: 59; 2-3yo: 54; 4-5yo: 58). 4.3. results. data analyses were conducted using r (r core team, 2013), using the package lme4 (bates et al. 2014a, 2014b). accuracy for controls was high (91.2%) (exp1d: mot: 96.1%; 2-3yo: 88.1%; 4-5yo: 90.8%; exp1f: mot: 96.0%; 2-3yo: 87.6%; 4-5yo: 91.8%). figure 2 shows the mean accuracy by force and age group after participant exclusion. on average, each context was seen by 38.9 participants, ranging between 19 and 59 times. analysis. overall, participants were accurate at recovering the original modal force (table 2). binomial tests (reported in appendix) show they differ from chance in all conditions. to test for the effect of force, we ran generalized linear mixed effects models, built with a maximal random effect structure allowed by our experimental design, testing accuracy (dependent variable, binomial), with force as fixed proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 136 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ effect and subject and item as random factors, and compare them with reduced models without force as a fixed effect (following barr et al., 2013) (glmer syntax of the full model: accuracy ~ force + (force|subject) + (1|item), reduced model: accuracy ~ 1 + (force|subject) + (1|item)).9 effect of force. for mother’s production, we find no difference in accuracy between possibility and necessity contexts (both overall and by experiment) (general mean accuracy: poss: 78%; nec: 77%). for child production, in both age groups, we find higher performance on possibility than necessity contexts (both overall and by experiment) (2-3yo: poss: 75%; nec: 60%; 4-5yo: poss: 82%; nec: 64%). results of the models’ comparison are given in table 3. effect of age. comparing child groups to mother (glmer syntax of the full model: accuracy ~ age + (1|subject) + (1|item); reduced model: accuracy ~ 1 + (1|subject) + (1|item)), we find no difference between children and adults for possibility modals (for both age groups), but the difference is significant for children necessity modals in exp1f (almost significant in 1d), and significant overall for 4 to 5-year-olds. there is no significant difference when comparing 2to 3-year-olds with 4to 5year-olds. table 4 summarizes the results. we checked that there was no effect of modal lemma (comparing exp1d to exp1f for all groups; glmer syntax of the full model: accuracy ~ experiment + (1|subject) + (1|item) (results are reported in appendix). finally, looking at the interactions age*force (full model: accuracy ~ force * age + 1|subject + 1|item)), we find a significant difference comparing 4to 5-year-olds to mothers, indicating that the difference in accuracy between possibility and necessity modals for child productions is larger than for mothers productions. the difference is almost significant comparing 2to 3-year-olds with mothers (p=0.06). we find no effect between the two child age groups. results are summarized in table 5. french english exp1d (pouvoir/devoir) exp1f (pouvoir/falloir) root1 (can/haveto) root2 (can/must) mot 2-to-3 3-to-5 mot 2-to-3 3-to-5 mot 2-to-3 mot 2-to-3 figure 2: mean accuracy by force and age group on exp1d and exp1f, with english (dieuleveut et al 2022, condition root_p1 and root_p2) as comparison. mother 2-3 y-o 4-5 y-o poss nece poss nece poss nece exp1d 80.4% (0.048) 76.6% (0.068) 77.7% (0.049) 60.5% (0.074) 82.3% (0.036) 66.4% (0.057) exp1f 75.3% (0.055) 78.2% (0.055) 74.5% (0.058) 59.3% (0.053) 82.5% (0.036) 64.6% (0.054) all 77,9% (0,036) 77,4% (0,043) 76,1% (0,038) 59,8% (0,044) 82,4% (0,025) 65,5% (0,039) table 2: mean accuracy (se) by age and force, experiment 1d and 1f. accuracy corresponds to the mean accuracy (how good participants were to guess correctly the force of the modal given the context) across the 20 contexts initially extracted from the corpus for each condition of force and age. each participant saw only 10 contexts (10 possibility, 10 necessity), randomly picked within pcibex. on average, each context was seen 39 times, ranging between 19 and 59 times. 9 answers were coded as 1 if the response was accurate, and 0 otherwise. the same procedure based on model comparisons was used for all subsequent experiments, so we don’t systematically report the reduced model. proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 137 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ mother 2-3 y-o 4-5 yo exp1d χ2(1)= 0.51, p = 0.47 (ns) χ2(1) = 4.05, p = 0.04 * χ2(1) = 5.28, p = 0.02 * exp1f χ2(1)= 0.26, p = 0.61 (ns) χ2(1) = 4.1, p = 0.04 * χ2(1) = 6.52, p = 0.01 * all χ2(1)= 0.02, p = 0.88 (ns) χ2(1)= 8.16, p = 0.004 ** χ2(1)= 11.6, p < .001 *** table 3: results of the model testing effect of force (possibility vs necessity), by age and lemma, experiment 1d and 1f. mot vs 2-3yo mot vs 4-5 yo 2-3yo vs 4-5yo possibility necessity possibility necessity possibility necessity exp1d χ2(1)= 0.21, p = 0.643 (ns) χ2(1)= 2.21, p = 0.14 (ns) χ2(1)= 0.01, p = 0.91 (ns) χ2(1)= 1.67, p = 0.20 (ns) χ2(1)= 0.17, p = 0.68 (ns) χ2(1)= 0.2, p = 0.66 (ns) exp1f χ2(1)= 0, p = 0 (ns) χ2(1)= 6.88, p = 0.01** χ2(1)= 0.67, p = 0.41 (ns) χ2(1)= 3.71, p = 0.054 (ns) χ2(1)= 0.63, p = 0.43 (ns) χ2(1)= 0.61, p = 0.43 (ns) all χ2(1)= 0.10, p = 0.75 (ns) χ2(1)= 8.21, p = 0.004 ** χ2(1)= 0.24, p = 0.63 (ns) χ2(1)= 5.12, p = 0.02 * χ2(1)= 0.72, p = 0.40 (ns) χ2(1)= 0.77, p = 0.38 (ns) table 4: results of the model testing effect of age (adult vs. child usage) experiment 1d and 1f. mother vs 2-3yo mother vs 4-5 yo 2-3 yo vs 4-5 yo exp1d χ2(1) = 0.67, p = 0.41 χ2(1) = 0.76, p = 0.38 χ2(1) = 0, p = 0.98 exp1f χ2(1) = 3.43, p = 0.06 χ2(1) = 3.84, p = 0.05* χ2(1) = 0.01, p = 0.92 all χ2(1) = 3.54, p = 0.06 χ2(1) =3.95, p = 0.047 * χ2(1) = 0, p = 0.97 table 5: results of the model testing interactions force * age, experiment 1d and 1f. 4.4. discussion. for the baseline on mothers’ productions, we find no effect of force: participants are accurate at guessing force based on the dialogue, with no difference between possibility and necessity contexts. but in children’s productions, we find one, always in the direction of children’s necessity modals being harder to identify (table 3). this indicates that children tend to “over-use” necessity modals: they use them when adults would rather use possibility modals. we therefore replicate results for english, and further, we show that they are still present among 4and 5-yearolds (see figure 2). examples (6) and (7), which led to particularly low accuracy rates, illustrate some children’s non-adult-like uses of necessity modals. (6) […] maman : et alors tu y arrivais bien ? mother: so you were good at it? enfant : oui ! child: yes! maman : c'est ce que t’as fait aujourd'hui ? mum: is that what you did today? enfant : oui ! child: yes! maman : ah ! mum: oh! enfant : et puis [il faut/on peut] attraper des papillons dans... dans un filet à papillons. child: and then [you have to/you can] catch butterflies in a butterfly net]. paris corpus, madeleine, 30028; mean accuracy: 17.4% (nobs = 23) (7) […] enfant : oh là c'est encore brûlant child: oh, it's still hot. maman : oh ça va. mum: oh, that's all right. enfant : non là child: no! enfant : arrête ! child: stop! maman : voilà ! mum: that's it! maman : on s'essuie les mains. mum: wipe your hands. enfant : et là aussi je [peux/dois] lécher. child: and i [can/have to] lick there too. lyon corpus, marie, 40005b; mean accuracy: 12.5% (nobs = 32) proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 138 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 5. experiment 2. to see whether participants’ lower performance on children’s necessity modals could come from expectations they might have for children to use more possibility modals rather than children’s effective misuses, we ran a follow-up that switches the roles of child and mother, i.e., attributing children’s utterances to mothers (chi-to-mot) and vice versa (mot-to-chi). this manipulation allows us to test for the effect of speaker. specifically, if the lower performance for children’s necessity modals in experiment 1 comes from participants expecting children to use more possibility modals (which could make sense, since possibility modals are indeed more frequent in children’s productions), we should obtain higher performance (reflecting more necessity guesses) if we attribute the same utterances to their mother. 5.1. methods. from the participants’ perspective, experiment 2 was identical to experiment 1.10 after instructions and a short training (3 examples on a/the), they had 40 dialogues to judge, presented in randomized order (20 controls based on tense and 20 trials: 10 possibility, 10 necessity, randomly selected from the list). we thus had four groups (exp2d: pouvoir vs devoir; exp2f: pouvoir vs falloir; mother vs child utterances). half of the trials (5/10 possibility, 5/10 necessity, 10/20 controls) tested the original dialogue (allowing us to replicate experiment 1’s results), whereas the other half had the speakers switched. figure 3 illustrates the manipulation. child child original utterance (experiment 1) chi-to-mot child utterance attributed to mother enfant : ... t'en laisses un petit coup. maman : merci. enfant : voilà. maman : merci. enfant : arrête d'aller là avec le ptit chevaux enfant : vous arrêtez d'aller là. enfant : parce que c'est après. enfant : qu'on ________ aller après. doit peut maman : ... t'en laisses un petit coup. enfant : merci. maman : voilà. enfant : merci. maman : arrête d'aller là avec le ptit chevaux maman : vous arrêtez d'aller là. maman : parce que c'est après. maman : qu'on ________ aller après. doit peut figure 3: example trials without/with the switch. 5.2. material. context selection. we had to exclude contexts when role switching resulted in infelicitous dialogues (like in (8) and (9)), or when the child or mother’s name was explicitly mentioned. to select the contexts, we asked three french speakers (including one of the authors) to rate the contexts used in experiment 1 on a scale between 1 and 4 (‘does the dialogue sound normal or strange?’; 1: normal; 4: weird) and kept contexts of a mean score < 3. this allowed us to keep 47% of the original lists (out of 177 contexts in experiment 1, we kept 37/80 originally from mot (46%) and 36/77 from chi (47%)). aside from the speaker’s label switch, we made no other change except agreement in two control contexts (e.g., “je vais aller” vs “je suis allée”). (8) […] enfant : ce que tu peux faire c'est mettre les toilettes dedans. ‘child: what you can do, is put toilets inside’ enfant : comme ça. ‘child: like that’ enfant: hopoo. ‘child: hopoo’ maman : je _____ faire pipi. ‘mother: i have to pee’ (paris, madeleine, 20206) 10 we tested only 2to 3-year-olds productions. experiment 1 on 4to 5-year-old was actually run afterwards. proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 139 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (9) […] maman : je veux manger dans ma chaise. ‘mother: i want to eat in my chair’ enfant : tu sais aller dans la chaise toute seule ? ‘child: you know how to go in your chair by your own?’ maman : c'est toi tu _____ mettre dans ma chaise. ‘mother: it’s you you have to put me in my chair’ (paris, julie, 20122) 5.3. participants. 120 french participants who had not taken part in experiment 1 were recruited on prolific (30 per group) (66 m, 49 f, 2 nb, 3 unknown; mean age: 32.5yrs). accuracy on controls was high (93.2%, with no difference between switch/noswitch control contexts: mot: 95.2%; mot-to-chi: 94.3%; chi: 91.7%; chi-to-mot: 91.8%). we excluded 2 participants due to low accuracy on controls (<75%). 5.4. results. table 6 summarizes mean accuracy in each condition. the first two rows provide results from experiment 1, first on all contexts, second after the context selection. the two last rows give results from experiment 2, first without the switch, then with the switch. we tested the effect of the switch using generalized linear mixed effects models, built with a maximal random effect structure allowed by our experimental design (details of the model are reported in appendix). we find no effect except for mother’s possibility modals, which overall and in exp2f lead to lower accuracy with the switch, indicating that participants tend to use more necessity modals when the modal is attributed to the child. this is the opposite of what we would obtain if participants were expecting children to use more possibility modals. table 6: mean accuracy (se) by age and force, experiment 2d and 2f (n=118), compared to experiment 1d/1f.on average, each context was seen 41 times in experiment 2 (20.9 times with original speaker, 20.2 times with switched speakers), ranging between 16 and 60 times. experiment 2d experiment 2f mother 2-3 yo mother 2-3 yo poss nece poss nece poss nece poss nece figure 5: effect of role switch, experiment 2d and 2f. expd (pouvoir vs devoir) expf (pouvoir vs falloir) mother 2-3 yo mother 2-3 yo poss nece poss nece poss nece poss nece i exp1 (all contexts) 80% 77% 78% 61% 75% 78% 75% 59% ii exp1 (kept for exp2) 80% 72% 80% 66% 73% 79% 78% 57% iii exp2 (no switch) 79,4% (0,054) 71,4% (0,099) 81,8% (0,061) 64,8% (0,105) 70,0% (0,095) 84,8% (0,039) 75,7% (0,088) 59,0% (0,095) iv exp2 (switch) 73,7% (0,074) 64,4% (0,109) 83,7% (0,048) 65,2% (0,094) 62,4% (0,085) 83,0% (0,028) 79,2% (0,068) 52,8% (0,102) proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 140 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ mot 2-3 yo possibility necessity possibility necessity exp2d χ2(1)=0.15, p=0.70 (ns) χ2(1)=2.28, p=0.13 (ns) χ2(1)=0.03, p=0.85 (ns) χ2(1)=0.04, p=0.83 (ns) exp2f χ2(1)=5.21, p=0.02 * χ2(1)=0.32, p=0.57 χ2(1)=0, p=1 (ns) χ2(1)=1.34, p=0.25 (ns) all χ2(1)=4.15, p=0.04 * χ2(1)=0.55, p=0.46 (ns) χ2(1)=0.01, p=0.93 (ns) χ2(1)=0.73, p=0.39 (ns) all χ2(1)=3.21, p=0.073 (ns) table 7: results of the model testing effect of switch (no switch vs switch), by condition and age group, experiment 2d and 2f. 6. discussion. our study makes three important points. first, we have shown that children’s difficulties with necessity modals are not limited to learners of english. we replicated dieuleveut el al. (2022)’s findings for english in french: children master possibility modals early, but use necessity modals them later on, less frequently, and, crucially, use them in a non-adult-like way: they “overuse” them. second, we have shown that these difficulties persist with older children (a point not assessed in dieuleveut et al. 2022 on english): our experiment shows that 4to 5-year-olds still use their necessity modals in a non-adult-like way. third, our study shows that the “delay” for necessity modals cannot be due only to low quantity of exposure. french children actually hear more necessity than possibility modals in their input: hearing more necessity modals doesn’t help. are children confused about the meaning of necessity modals? or, is it simply that they do not know yet in which contexts they are appropriate? more research will be needed to answer these questions. children might also fail to use necessity modals as adults do because of problems determining the modal’s domain of quantification (specifically, consider a smaller set of worlds than adults). politeness could also be involved: adults tend to use possibility modals to express requests, in contexts in which children might more directly use necessity modals. one of our goal is to tease these hypotheses apart, acknowledging that they are not necessarily mutually exclusive. taken together, our results call for extension to other logical scales, like some/all or sometimes/always, where similarly subset problems arise—do children have similar difficulty with all and always? the method we used, which uses adult’s judgments to evaluate children’s spontaneous productions, based on existing corpus data, has several advantages: it is quite easy to deploy, simple from the participant’s perspective, and therefore easily generalizable to other languages and to other cases of word learning. references acredolo, curt & karen horobin. 1987. development of relational reasoning and avoidance of premature closure. developmental psychology 23(1): 13–21. https://doi.org/10.1037/00121649.23.1.13. barner, david, neon brooks & alan bale. 2011. accessing the unsaid: the role of scalar alternatives in children’s pragmatic inference. cognition 118(1): 84–93. https://doi.org/10.1016/j.cognition.2010.10.010. barr, dale j., roger levy, christoph scheepers & harry j. tily. 2013. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language 68(3): 255–278. https://doi.org/10.1016/j.jml.2012.11.001. bates, douglas, martin mächler, benjamin m. bolker & steven c. walker. 2014. fitting linear mixed-effects models using lme4. arxiv preprint arxiv:1406.5823. https://doi.org/10.18637/jss.v067.i01 berwick, robert c. 1985. the acquisition of syntactic knowledge. cambridge: mit press. proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 141 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ chierchia, gennaro, stephen crain, maria teresa guasti, andrea gualmini & luisa meroni. 2001. the acquisition of disjunction: evidence for a grammatical view of scalar implicatures. in proceedings of the 25th boston university conference on language development, eds. a.h.-j. do, l. domínguez, and a. johansen, 157–168. somerville: cascadilla press. chu, xiaoquan. 2008. les verbes modaux du français. editions ophrys. cournane, ailís & sandrine tailleur. 2020. la production épistémique chez l’enfant francophone: complexité syntaxique et ordre d’acquisition. arborescences : revue d’études françaises, 10, 47–72. cournane, ailís. 2015b. modal development: input-divergent l1 acquisition in the direction of diachronic reanalysis, phd dissertation, university of toronto. cournane, ailís. 2021. revisiting the epistemic gap: evidence for a grammatical source. language acquisition 28(3): 215-240. https://doi.org/10.1080/10489223.2020.1860054. crain, stephen & rosalind thornton. 1998. investigations in universal grammar. cambridge: mit press. deal, amy rose. 2011. modals without scales. language 87(3): 559–585. https://doi.org/10.1353/lan.2011.0060 demuth, katherine & annie tremblay. "prosodically-conditioned variability in children's production of french determiners." journal of child language 35.1 (2008): 99-127. dieuleveut, anouk, annemarie van dooren, ailís cournane & valentine hacquard. 2019a. learning modal force: evidence from children’s production and input. in proceedings of the 2019 amsterdam colloquium, 111–122. dieuleveut, anouk, annemarie van dooren, ailís cournane & valentine hacquard. 2019b. acquiring the force of modals: sig you guess what sig means? in proceedings of the 43rd annual to the boston university conference on language development (bucld43), ed. megan m. brown and brady dailey, 189-202. dieuleveut, anouk. 2021. finding modal force, phd dissertation, university of maryland. dieuleveut, anouk. 2023. je peux, ou je dois? faudrait savoir! acquiring modal force: evidence from french. in proceedings of the 47rd annual to the boston university conference on language development (bucld47). gillette, jane, henry gleitman, lila gleitman & anne lederer. 1999. human simulations of vocabulary learning. cognition 73(2): 135–176. https://doi.org/10.1016/s0010-0277(99)000360. gleitman, lila r., kimberly cassidy, rebecca nappa, anna papafragou & trueswell, john c. 2005. hard words. language learning and development, 1(1), 23-64. grice, herbert paul. 1975. logic and conversation. in speech acts, 41–58, ed. j.p. kimball. syntax and semantics vol. 3. leiden: brill. https://doi.org/10.1163/9789004368811_003 gualmini, andrea & bernhard schwarz. 2009. solving learnability problems in the acquisition of semantics. journal of semantics 26(2): 185–215. https://doi.org/10.1093/jos/ffp002. hacquard, valentine. 2011. modality. in semantics: an international handbook of natural language meaning, eds. c. maienborn, k. von heusinger, and p. portner, 1484–1515. berlin: mouton de gruyter. horn, laurence r. 1972. on the semantic properties of logical operators, phd dissertation, ucla. iatridou, sabine & hedde zeijlstra. 2013. negation, polarity, and deontic modals. linguistic inquiry 44(4): 529–568. http://dx.doi.org/10.1162/ling_a_00138. proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 142 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ jeretič, paloma. 2018. evidence for children’s dispreference for weakness: a corpus study. manuscript, new york university. kratzer, angelika. 1981. partition and revision: the semantics of counterfactuals. journal of philosophical logic 10(2): 201–216. https://doi.org/10.1007/bf00248849. kratzer, angelika. 1991. modality. semantics: an international handbook of contemporary research, eds. c. maienborn, k. von heusinger, and p. portner, 639–650. berlin: mouton de gruyter. macwhinney, brian. 2000. the childes project: the database. vol. 2. hove, uk: psychology press. manzini, m. rita & kenneth wexler. 1987. parameters, binding theory, and learnability. linguistic inquiry 18(3): 413–444. https://www.jstor.org/stable/4178549. morgenstern, aliyah, and christophe parisse. "the paris corpus." journal of french language studies 22.1 (2012): 7-12. moscati, vincenzo, likan zhan & peng zhou. 2017. children’s on-line processing of epistemic modals. journal of child language 44(5): 1025–1040. https://doi.org/10.1017/s0305000916000313. noveck, ira a. 2001. when children are more logical than adults: experimental investigations of scalar implicature. cognition 78(2): 165–188. https://doi.org/10.1016/s0010-0277(00)001141. ozturk, ozge & anna papafragou. 2015. the acquisition of epistemic modality: from semantic meaning to pragmatic interpretation. language learning and development 11(3): 191–214. https://doi.org/10.1080/15475441.2014.905169. papafragou, anna. 1998. the acquisition of modality: implications for theories of semantic representation. mind & language 13(3): 370–399. https://doi.org/10.1111/1468-0017.00082. piantadosi, steven t. 2011. learning and the language of thought, phd dissertation, mit. r core team. 2013. r: a language and environment for statistical computing. https://www.r-project.org/. rasin, ezer & athulya aravind. 2020. the nature of the semantic stimulus: the acquisition of every as a case study. natural language semantics 29: 339–375. https://doi.org/10.1007/s11050-020-09168-6. skordos, dimitrios & anna papafragou. 2016. children’s derivation of scalar implicatures: alternatives and relevance. cognition 153: 6–18. https://doi.org/10.1016/j.cognition.2016.04.006 theakston, anna l., elena v. lieven, julian m. pine & caroline f. rowland. 2001. the role of performance limitations in the acquisition of verb-argument structure: an alternative account. journal of child language 28(1): 127–152. van dooren, annemarie, anouk dieuleveut, ailís cournane & valentine hacquard. 2022. figuring out root and epistemic uses of modals: the role of the input. journal of semantics. xu, fei & joshua b. tenenbaum. 2007. word learning as bayesian inference. psychological review 114(2): 245–272. https://doi.org/10.1037/0033-295x.114.2.245. yanovich, igor. 2016. old english *motan, variable-force modality, and the presupposition of inevitable actualization. language 92(3): 489–521. zehr, jeremy & florian schwarz. (2018). penncontroller for internet based experiments (ibex). https://doi.org/10.17605/osf.io/md832 proceedings of elm 3: 130-143, 2025 anouk dieuleveut and ira noveck: devoir, ou pouvoir, that is the question. 143 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ linguistic and social meaning match: an experiment on modal concord in english mingya liu & stephanie rotter* abstract. modal concord (mc) refers to the phenomenon where two modal elements of the same flavor and force in a sentence yield an interpretation of single modality (sm). in this paper, we report on an experimental study on mc in english, addressing their linguistic and social meaning. our results show a strengthening effect by necessity mc and a weakening effect of possibility mc in that significantly higher speaker commitment ratings were received for necessity mc vs. sm constructions (i.e., must certainly vs. must) with the reverse pattern for possibility modal constructions (i.e., may possibly vs. may). furthermore, mc and sm were shown to differ in social meanings, suggesting a correlation between the meaning strength of a linguistic expression and the social perception of the speaker. keywords. modal concord; social meaning; experiment; english 1. introduction. modality is one of the key empirical areas in linguistics, and also one of the most complex ones. this paper deals with the phenomenon ‘modal concord’ (mc) and addresses the linguistic and social meaning of mc constructions. as shown in (1), a sentence can be unmodalized (1-a), contain a single modal (sm) element such as a modal auxiliary or adverb (1-b), or include two modal elements including a modal auxiliary and an adverb (1-c). (1) a. it is the case. b. it may be the case. / it is possibly the case. (sm) c. it may possibly be the case. (mc) in (1-c), the two modal elements are both weak epistemic modals and their co-occurrence can yield the same meaning as just one of them—this phenomenon is labeled as mc in the literature (geurts & huitink 2006, zeijstra 2007, yalcin 2007), with the discussion dating back to at least the 70s, see an earlier quote and a more recent one below. “in most dialects of english not more than one modal verb can occur within the same clause. but both a modal verb and a modal adverb may be combined. when this happens a distinction is to be drawn between modally harmonic∗ and modally nonharmonic∗ combinations. for example, ‘possibly’ and ‘may’, if each is being used epistemically, are harmonic, in that they both express the same degree of modality, whereas ‘certainly’ and ‘may’ are, in this sense, modally non-harmonic. it has been pointed out by halliday (1970a: 331) that the adverb and the modal verb may, and normally do, “reinforce each other” in a modally harmonic combination; so that, ... or there is a kind of concord running through the clause, which results in the double realization of a single modality.” (lyons 1977: 807-808). *this work was funded by the deutsche forschungsgemeinschaft (dfg, german research foundation) – sfb 1412, 416591334. authors: mingya liu, humboldt-universität zu berlin (mingya.liu@hu-berlin.de) & stephanie rotter, humboldt-universität zu berlin (rotterst@hu-berlin.de). proceedings of elm 3: 224-235, 2025 c©2025 mingya liu and stephanie rotter published by the lsa with permission of the author(s) under a cc by license. 224 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ “if two modal elements are of the same modal type (epistemic/ deontic/...) and have similar quantificational force (universal/existential), the most salient reading is mostly not a cumulative one but a concord reading, where the semantics seems to contain only one modal operator.” (zeijstra 2007: 137) with the ‘meaning-equivalence’ assumption for mc and sm, we can make the inference that the optional adverb possibly in (1-c) is semantically vacuous. this is debatable. for example, some authors hold the view that multiple modal expressions have one semantic role but the adverb does make a semantic contribution, giving rise to speaker bias (i.e., weakened or strengthened speaker commitment to the modified proposition), see the discussion of ‘modal spread’ by giannakidou & mari (2018). despite the numerous attempts to integrate mc into semantic theories, the empirical picture remains unclear: do mc constructions really share the same semantics as their sm counterparts? another question that arises naturally is the use of mc in comparison to sm: assuming mc and sm are ‘functional equivalents’, that is, variants of a linguistic variable with the same semantic interpretation (labov 2006: 145-146), why would a speaker choose to use mc over sm? this question is specific to mc in the modal domain, as it is available across languages and varieties of english, in contrast to the phenomenon of ‘multiple modals’ (e.g., might could), which is much more restricted, that is, only possible in some varieties of english (see the first sentence in lyons’ 1977 quote above, and kortmann & schneider 2004 on colloquial american and appalachian american english) and in general, in a smaller set of natural languages. based on the sociolinguistic literature and more recent work on register in lüdeling et al. (2022) and pescuma et al. (2023), we assume that the choice between mc and sm can be register-driven, that is, the speaker chooses one or the other depending on the properties of the situational contexts with regard to, for example, interlocutor relations or communicative purposes. furthermore, the choice can be better understood in relation to the social meaning, that is, “the set of inferences that can be drawn on the basis of how language is used in a specific interaction” (hall-lew et al. 2021; 3). for example, glass (2015) reports on a social meaning study of three different semi-modals of universal force in english, see (2), and argues that “you need to has a slightly different semantics than you have/got to, and that this subtlety gives it a unique social meaning that depends on whether the speaker is licensed to tell the hearer what’s good for him” (p.87). similarly to that study, our study also addresses the social meaning of modal expressions of the same—existential or universal—force: more specifically, when a speaker utters mc in a specific context, what pragmatic inferences can be drawn about the speaker’s backgrounds and personae? (2) you {need to /have to/gotta} admire her for persevering. (glass 2015: example (1)) in the following, we report on an experimental study on mc in english, which is guided by the following questions: (rq1) how are mc constructions interpreted? we tested mc and sm constructions, using speaker commitment ratings, and found that their interpretations differ. (rq2) how are mc constructions perceived? we tested the social meaning of mc and sm, as well as that of possibility and necessity modals, and found that their social meanings differ. (rq3) how does the perception of doubling constructions differ? we will discuss the results of mc in the current study with the results of negative concord (nc) in our previous work (rotter & liu 2025, proceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 225 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 2024) and show that the social meanings of mc and nc differ. the paper is structured as follows: section 2 presents the method and the results of the experiment. section 3 discusses the results, the limitations and future outlooks of the study, and concludes the paper. 2. experiment. 2.1. material and design. the experiment was implemented with pcibex farm and hosted on its platform (https://farm.pcibex.net/). we used a 2x2-factorial design with the factors force (possibility vs. necessity) and number (of used modal elements, i.e., mc vs. sm) in a latin square design with four lists. we used 12 critical, see (3), and 29 filler items of similar structure. (s1) remained the same across all items of the study, see (3). (s2) contained the force and number manipulation. (q1) was used to measure the interpretation of the sentence in terms of speaker commitment ratings (liu et al. 2021) and (q2) to measure the grammaticality ratings. (3) (s1) somebody says: (s2) “i / may possiblymc / have gotten / the wrong address.” (possibility) “i / maysm / have gotten / the wrong address.” (possibility) “i / must certainlymc / have gotten / the wrong address.” (necessity) “i / mustsm / have gotten / the wrong address.” (necessity) (q1) does the person believe they have gotten the wrong address? (q2) is the sentence grammatical? 2.2. procedure. the experiment consisted of four parts: (p1) consent, (p2) two practice sentences to get familiarized with the experimental set-up, (p3) main study, and (p4) demographic and language background survey. participants were instructed to use the space bar to continue. each trial in (p2/3) started with a fixation cross in the middle of the screen. then (s1) appeared on the screen. (s2) was presented as a self-paced reading task; the length and position of each chunk was indicated by a line on the screen. previous chunks disappeared so that only the current chunk was visible. finally, the entire sentence (s2) appeared on the screen together with the following social meaning measures: social backgrounds of the speaker (low/high socioeconomic status and low/high education), and persona of the speaker (in/formal, un/confident, im/polite, un/friendly, cold/warm, un/cool, and obedient/rebellious); the measures appeared in groups of three and within each group, the order was randomized. in the two last screens, the participants first rated the speaker commitment, that is, interpretation (q1-certainly no/certainly yes), and then the grammaticality of the sentence (q2-certainly no/certainly yes). we chose this order to allow for unbiased speaker commitment ratings. all the measures and questions used a 7-point likert scale in which the endand midpoints were labels (e.g., 1: ungrammatical – 4: undecided – 7: grammatical). participants used the mouse click to provide their answer. 2.3. participants. we collected data from 104 native speaker of us english on the crowdsourcing platform prolific (https://www.prolific.co/). the experiment took roughly 35 minutes and participants received monetary compensation. all participants provided their informed consent as approved by the ethics committee of the deutsche gesellschaft für sprachproceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 226 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ wissenschaft (dgfs) in the context of sfb 1412 ‘register’. based on the inclusion criteria (see 2.4 for details), we analyzed the data of 93 participants from 37 of the 50 us states (mean age=39.2/sd=10.9, range=[19, 63]; female n=48, male n=43, non-binary n=2). 2.4. data analysis. the data was processed and analyzed with the open source software ‘r’ (version 4.1.1, r core team 2024) in the rstudio environment (rstudio team 2024). we defined demographic (native english speaker and aged between 18 to 65 years) and attendance (75% of the chunks were attended for at least 150msec) inclusion criteria: first, all participants matched the demographic criteria. second, the data of 11 participants were excluded based on the attendance. furthermore, we removed the data of 254 (10.2%) trails, in which at least one chunk was attended for less than 150msec. we calculated separate models for each measure with the remaining data in the cumulative link function model framework (howcroft & rieser 2021, liddell & kruschke 2018) using the package ‘ordinal’ (christensen 2019). the link function of each model was determined as the one with the highest log-likelihood value among the five possible link functions (i.e., probit, logit, cauchit, loglog, and cloglog) (christensen 2019). each model contained the sum-coded factors force (possibility: 0.5, necessity: −0.5) and number (mc: 0.5, sm: −0.5), as well as their two-way interaction force×number (2-int). if the interaction turned out significant, we conducted a sub-analysis by splitting the data along the significant main factor. the sub-analysis included the same coding of the main effects. we used the most parsimonious model approach to obtain the random effect structures; the used model structures are indicated in the respective result section. log-likelihood ratio test comparisons of nested models (bates et al. 2018) were used to obtain p-values, which were defined as significant at a value below 0.05. we reduced the complexity of models if their null-model did not converge. we used the model with the smaller akaike information criterion (aic) if the model choice was between subject or item intercepts. all statistical values of means, estimates and the like are rounded to the second decimals except for p-values smaller than 0.01. 3. results. below, we first present the results of the interpretation and grammaticality measures and then the results of the social meaning measures. 3.1. interpretation and grammaticality. the output of the models using interpretation (q1) and grammaticality (q2) ratings as dependent variable are shown in table 2. for the interpretation ratings, the logit-link function model including random subject intercepts and slopes for force, number, and their interaction fit the data best. the results showed a significant main effect of force in that possibility conditions were rated lower than necessity conditions (β̂=−3.65, χ2(1)=95.99, p<0.001). there was a main effect of number in that mc was rated higher than sm (β̂=0.60, χ2(1)=15.91, p<0.001). furthermore, the 2-way interaction turned out significant (β̂=−1.85, χ2(1)=41.51, p=<0.001), which results from a cross-over effect: • weakening effect in possibility modals: mc received significantly lower speaker commitment ratings than sm (β̂=−0.35, χ2(1)=6.79, p<0.009) • strengthening effect in necessity modals: mc received significantly higher speaker commitment ratings than sm (β̂=1.50, χ2(1)=36.92, p<0.001). proceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 227 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: by-subject means (transparent dots) of the rating measures in comparison to overall means (opaque dots with error bars). panel a depicts the ratings of the interpretation and grammaticality, panel b those of the social background measures, and panel c those of the the persona measures. the x-axis indicates the specific measures. the top scale shows the ratings from the necessity conditions (abbreviated as nece), the one below those of the possibility conditions (abbreviated as poss). the colors indicate the factor number with mc (i.e., modal concord) in blue and sm (i.e., single modal) in yellow. the y-axis depicts the ratings on a 7-point likert scale. ses abbreviates socioeconomic status. proceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 228 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ n med mean sd se med mean sd se force number interpretation grammaticality poss mc 542 5 5.11 0.96 0.04 6 5.06 2.05 0.09 poss sm 541 5 5.22 0.98 0.04 7 6.45 1.09 0.05 nece mc 536 7 6.40 0.90 0.04 6 4.90 2.18 0.09 nece sm 539 6 6.12 0.85 0.04 7 6.41 1.16 0.05 socioeconomic status education poss mc 542 5 4.73 1.30 0.06 5 4.69 1.37 0.06 poss sm 541 5 4.87 1.11 0.05 5 4.94 1.12 0.05 nece mc 536 5 4.85 1.40 0.06 5 4.80 1.46 0.06 nece sm 539 5 4.87 1.18 0.05 5 4.89 1.19 0.05 formality poss mc 542 5 4.73 1.61 0.07 poss sm 541 5 4.84 1.38 0.06 nece mc 536 5 4.99 1.66 0.07 nece sm 539 5 4.76 1.40 0.06 politeness confidence poss mc 542 5 5.36 1.17 0.05 4 4.03 1.61 0.07 poss sm 541 5 5.45 1.11 0.05 4 4.36 1.51 0.06 nece mc 536 5 5.39 1.16 0.05 6 5.48 1.41 0.06 nece sm 539 5 5.28 1.15 0.05 5 5.19 1.39 0.06 friendliness warmth poss mc 542 5 4.92 1.17 0.05 5 4.82 1.19 0.05 poss sm 541 5 5.03 1.14 0.05 5 4.94 1.16 0.05 nece mc 536 5 4.76 1.21 0.05 4 4.62 1.25 0.05 nece sm 539 5 4.97 1.15 0.05 5 4.86 1.17 0.05 coolness rebelliousness poss mc 542 4 4.29 1.33 0.06 3 3.11 1.33 0.06 poss sm 541 4 4.56 1.27 0.05 3 3.09 1.34 0.06 nece mc 536 4 4.20 1.32 0.06 3 3.04 1.34 0.06 nece sm 539 4 4.53 1.27 0.05 3 3.12 1.36 0.06 table 1: descriptive statistics of the interpretation (q1), grammaticality (q2), social background, and persona ratings. poss stands for possibility conditions, nece for necessity, mc for modal concord, and sm single modal. med abbreviates median, sd standard deviation, and se standard error. for the grammaticality ratings, the logit-link function model including random subject and item intercepts, both with slopes for force and number fit the data best. the results showed a significant main effect of number in that mc was rated less grammatical than sm (β̂=−3.48, χ2(1)=83.01, p<0.001). proceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 229 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ fixed effects model comparison fixed effects model comparison β̂ se χ2(1) p-value β̂ se χ2(1) p-value interpretation grammaticality force −3.65 0.30 95.99 <0.001⋆ 0.21 0.29 0.53 0.47 number 0.60 0.15 15.91 <0.001⋆ −3.48 0.36 83.01 <0.001⋆ 2-int −1.85 0.28 41.51 <0.001⋆ 0.15 0.23 0.43 0.51 table 2: output of the model using interpretation and grammaticality ratings. 2-int abbreviates the two way interaction force×number. se stands for the standard error. 3.2. social background. the output of the models using social background ratings as dependent variables are shown in table 3. fixed effects model comparison fixed effects model comparison β̂ se χ2(1) p-value β̂ se χ2(1) p-value socioeconomic status education level force −0.15 0.05 9.13 <0.003⋆ −0.11 0.05 4.84 < 0.03⋆ number −0.01 0.05 0.01 0.92 −0.06 0.05 1.25 0.26 2-int −0.26 0.10 6.58 0.01⋆ −0.31 0.10 9.62 <0.002⋆ table 3: output of the model using the social background ratings. 2-int abbreviates the two way interaction force×number. se stands for the standard error. for the ses and education level ratings, the loglog-link function model including random subject intercepts fit the data best. the results showed a significant main effect of force in that possibility conditions received lower ses (β̂=−0.15, χ2(1)=9.13, p<0.003) and education level (β̂=−0.11, χ2(1)=4.84, p<0.03) ratings than necessity conditions. furthermore, the 2-way interaction showed a significant effect in the ses model (ses-model: β̂=−0.26, χ2(1)=6.58, p=0.01; education level-model: β̂=−0.31, χ2(1)=9.62, p<0.002), which results from mc receiving lower ses (β̂=−0.18, χ2(1)=5.55, p<0.02) and education level (β̂=−0.26, χ2(1)=11.92, p<0.001) ratings than sm only in possibility conditions. 3.3. personae. the output of the models using personae ratings as dependent variables are shown in table 4. for the formality ratings, the cloglog-link function model including random subject intercepts with force, number, and their 2-way interaction as well as random item intercepts fit the data best. the results showed a significant main effect of force in that possibility conditions received lower formality ratings than necessity conditions (β̂=−0.18, χ2(1)=4.19, p=0.04). furthermore, the 2-way interaction showed a significant effect (β̂=−0.50, χ2(1)=9.12, p<0.003), which results from mc receiving higher formality ratings than sm only in necessity conditions (β̂=0.45, χ2(1)=8.80, p=0.003). for the politeness ratings, the cloglog-link function model including random subject intercepts fit the data best. the results showed no significant effects. for the confidence ratings, the logit-link function model including random subject and item intercepts fit the data best. the results showed a significant main effect of force in that possibilproceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 230 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ fixed effects model comparison fixed effects model comparison β̂ se χ2(1) p-value β̂ se χ2(1) p-value formality force −0.18 0.09 4.19 0.04⋆ number 0.21 0.12 3.17 0.08 2-int −0.50 0.16 9.12 <0.003⋆ politeness confidence force 0.08 0.06 2.06 0.15 −1.79 0.09 453.50 <0.001⋆ number 0.03 0.06 0.22 0.64 0.05 0.08 0.45 0.50 2-int −0.21 0.11 3.56 0.06 −1.03 0.16 41.28 <0.001⋆ friendliness warmth force 0.14 0.05 7.16 <0.008⋆ 0.16 0.05 9.47 0.002⋆ number −0.20 0.05 14.01 <0.001⋆ −0.26 0.05 23.74 <0.001⋆ 2-int 0.08 0.10 0.52 0.47 0.13 0.10 1.62 0.20 coolness rebelliousness force 0.02 0.05 0.15 0.70 0.02 0.05 0.10 0.76 number −0.42 0.05 62.42 <0.001⋆ −0.05 0.07 0.45 0.50 2-int 0.20 0.10 3.62 0.06 0.13 0.10 1.91 0.17 table 4: output of the model using the social background ratings. 2-int abbreviates the two way interaction force×number. se stands for the standard error. ity conditions received lower ratings than necessity conditions (β̂=−1.79, χ2(1)=453.50, p<0.001). furthermore, the 2-way interaction turned out significant (β̂=−1.03, χ2(1)=41.28, p<0.001), which results from a cross-over effect: mc received significantly lower confidence ratings in possibility conditions than sm (β̂=−0.66, χ2(1)=20.25, p<0.001), while mc received higher ratings than sm in necessity conditions (β̂=0.52, χ2(1)=16.76, p<0.001). for the friendliness and warmth ratings, the loglog-link model with random subject intercepts fit the data best. the results showed a significant main effect of force in that possibility conditions received higher friendliness (β̂=0.14, χ2(1)=7.16, p<0.008) and warmth (β̂=0.16, χ2(1)=9.47, p=0.002) ratings than necessity conditions. there was a significant main effect of number in that mc received lower friendliness (β̂=−0.20, χ2(1)=14.01, p<0.001) and warmth (β̂=−0.26, χ2(1)=23.74, p<0.001) ratings than sm conditions. for the coolness ratings, the loglog-link model with random subject intercepts fit the data best. the results showed a significant effect of number in that mc received lower coolness ratings than sm (β̂=−0.42, χ2(1)=62.42, p<0.001). for the rebelliousness ratings, the probit-link model with random subject intercepts and slopes for number fit the data best. the results showed no significant effects. 4. discussion and conclusion. the results of the study are summarized in table 5. first, mc was rated as less grammatical (with the mean ratings 5.06/4.90 for possibility/necessity mc) than sm. we take this to be in line with the more restricted distribution of mc. in relation to (rq1), our study shows that mc conveys a weakened speaker commitment proceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 231 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ force number 2-int force×number interpretation poss < ness mc > sm poss: mc < sm ness: mc > sm grammaticality – mc < sm – – socioeconomic status poss < ness – poss: mc < sm – education level poss < ness – poss: mc < sm – formality poss < ness – – ness: mc > sm politeness – – – – confidence poss < ness – poss: mc < sm ness: mc > sm friendliness poss > ness mc < sm – – warmth poss > ness mc < sm – – coolness – mc < sm – – rebelliousness – – – – table 5: summary of the results. the abbreviation ‘poss’ stands for possibility and ‘ness’ for necessity conditions. mc stands for modal concord and sm for single modal. 2-int abbreviates the two-way interaction force×number. the symbol < indicates that the left entity is smaller than the right one; > represents the reverse. the symbol – indicates the lack of a significant effect. than sm in possibility modals and a strengthened speaker commitment in necessity modals. this finding challenges the semantic equivalence assumption for mc and sm as well as the claim that modal adverbs in mc constructions have no semantic contributions. one way out would be to analyze weakening or strengthening as a pragmatic effect of, say, “reinforcing” (halliday 1970, see the quote above from lyons 1977) of the respective modal force (i.e., existential or universal): in simple words, doubling does not add extra contributions to the semantics but give rise to enriched meanings about speaker assumptions. alternatively, such effects can be taken as arguments for the ‘modal spread’ analysis of giannakidou & mari (2018), by which the weakening and strengthening effect can be analyzed as semantic contributions of modal adverbs. while we do not provide any formal analysis for this in the paper, we see several arguments in favor of the ‘modal spread’ analysis. firstly, if we take modal adverbs such as certainly in must certainly to have no semantics in mc constructions, we need to stipulate a different lexical entry for the same adverb in other contexts, for example, where it co-occurs with a modal verb of a different force. our preliminary corpus study shows that (4-a) is a naturally-occurring combination, which is “modally non-harmonic” (see the quote above from lyons 1977)—in contrast, the well-formedness of (4-b) is questionable, a topic we leave for future work. given the well-formedness of (4-a) as well as cases where modal adverbs occur without verbs, it strikes us as unattractive to analyze modal adverbs as semantically vacuous for mc constructions. (4) a. it may certainly be the case. b. ?/*it must possibly the base. in relation to (rq2), (i) mc was rated as less friendly, warm, and cool than sm. this is something which we neither predicted nor have a good explanation for, and thus needs to be further investigated in future studies. crucially, certain measures showed a significant number*force interaction in that (ii) in possibility modals, mc was rated as significantly lower than sm in ses, proceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 232 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ education and confidence levels and (iii) in necessity modals, mc was rated as more formal and more confident than sm. relating these results with the results of the interpretation, it seems that weaker statements give rise to more negative perceptions and stronger ones to more positive perceptions, indicating a positive correlation between meaning strength and social perception. in relation to (rq3), we only compare mc with nc here, see a brief grammatical comparison of these two doubling phenomena in zeijstra (2007). sociolinguistics and recent research on social meaning have both documented work on nc, which is more prevalent in different varieties of english and socially stigmatized (eckert 1989, 2019, moore 2021, rotter & liu 2025). in comparison, there is little work on the social meaning of mc, which unlike nc, has a wide distribution in english(es) and is not socially stigmatized. in recent work, rotter & liu (2025, 2024) conducted a series of studies on the register-sensitivity and the social meaning of nc compared to negative polarity items (npis) in american english, see (5). the findings (in the data by mostly self-reported non-dialect speakers of us english) are that (i) nc use is register-sensitive in that it was rated less appropriate in formal than in informal contexts, (ii) comprehenders interpreted nc similarly to npis, but (iii) associated nc use with lower socioeconomic and education levels, as well as less formal but more rebellious personae of the speaker. the current study, in combination with our previous work, provides experimental evidence that the doubling phenomena of mc and nc have very distinct social meanings. (5) a. i didn’t see nothing. (nc construction) b. i didn’t see anything. (npi construction) before we finish the paper, we would like to briefly address the limitations and future outlooks of our study. first, our study compared mc with sm constructions with a single modal auxiliary. the design can be extended to the comparison between mc, sm-verb and sm-adverb (e.g., must certainly vs. must vs. certainly and may possibly vs. may vs. possibly). secondly, we used isolated sentences in the study, which was our first step towards understanding the social meaning of mc, but it is evident that the set of inferences to be drawn from its use can differ greatly depending on the specific interaction. the next step would therefore be to incorporate context and further explore the question of whether mc is register-sensitive, and what social meanings arise in relation to its use in different situational contexts. data repository. the data set for this study can be found in the online repository: https: //osf.io/47rkg/. authorship contribution statement. mingya liu: conceptualization, formal analysis, methodology, writing original draft, supervision, project administration, funding acquisition. stephanie rotter: methodology, data collection, data analysis and visualization, writing original draft. references bates, douglas, reinhold kliegl, shravan vasishth & harald baayen. 2018. parsimonious mixed models. arxiv 1506. https://doi.org/10.48550/arxiv.1506.04967. christensen, rune haubo bojesen. 2019. ordinal: regression models for ordinal data cran. r-project.org/package=ordinal. r package version 2023.12-4. proceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 233 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ eckert, penelope. 1989. jocks and burnouts: social categories and identity in the high school. new york: teachers college press. eckert, penelope. 2019. the limits of meaning: social indexicality, variation, and the cline of interiority. language 95(4). 751–776. https://doi.org/10.1353/lan.0.0239. geurts, bart & janneke huitink. 2006. modal concord. in paul dekker & hedde zeijlstra (eds.), proceedings of the esslli 2006 workshop concord phenomena at the syntax-semantics interface, 15–20. malaga: university of malaga. giannakidou, anastasia & alda mari. 2018. the semantic roots of positive polarity: epistemic modal verbs and adverbs in english, greek and italian. linguistics and philosophy 41(6). 623–664. https://doi.org/10.1007/s10988-018-9235-1. glass, lelia. 2015. strong necessity modals: four socio-pragmatic corpus studies. in paul dekker & hedde zeijlstra (eds.), university of pennsylvania working papers in linguistics, vol. 21(2), 79–88. hall-lew, lauren, emma moore & robert j. podesva. 2021. social meaning and linguistic variation: theoretical foundations. in lauren hall-lew, emma moore & robert j. podesva (eds.), social meaning and linguistic variation, 1–24. cambridge: cambridge university press. https://doi.org/10.1017/9781108578684.001. halliday, michael a. k. 1970. functional diversity in language as seen from a consideration of mood and modality in english. foundations of language 6. 322–361. howcroft, david m. & verena rieser. 2021. what happens if you treat ordinal ratings as interval data? human evaluations in nlp are even more under-powered than you think. in marie-francine moens, xuanjing huang, lucia specia & scott wen-tau yih (eds.), proceedings of the 2021 conference on empirical methods in natural language processing, 8932– 8939. online and punta cana, dominican republic: association for computational linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.703. https://aclanthology. org/2021.emnlp-main.703. accessed 2022-01-03. kortmann, bernd & edgar w. schneider. 2004. a multimedia reference tool. volume 1: phonology. volume 2: morphology and syntax. berlin, boston: de gruyter mouton. https://doi.org/10.1515/9783110197181. labov, william. 2006. the social stratification of english in new york city. cambridge: cambridge university press 2nd edn. https://doi.org/10.1017/cbo9780511618208. liddell, torrin m. & john k. kruschke. 2018. analyzing ordinal data with metric models: what could possibly go wrong? journal of experimental social psychology 79. 328–348. https://doi.org/10.1016/j.jesp.2018.08.009. liu, mingya, stephanie rotter & anasatasia giannakidou. 2021. bias and modality in conditionals: experimental evidence and theoretical implications. journal of psycholinguistic research 50. 1369–1399. https://doi.org/10.1007/s10936-021-09813-z. lyons, john. 1977. semantics, vol. 2. cambridge: cambridge university press. https://doi.org/10.1017/cbo9780511620614. lüdeling, anke, artemis alexiadou, aria adli, karin donhauser, malte dreyer, markus egg, anna helene feulner, natalia gagarina, wolfgang hock, stefanie jannedy, frank kammerzell, pia knoeferle, thomas krause, manfred krifka, silvia kutscher, beate lütke, proceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 234 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ thomas mcfadden, roland meyer, christine mooshammer, stefan müller, katja maquate, muriel norde, uli sauerland, stephanie solt, luka szucsich, elisabeth verhoeven, richard waltereit, anne wolfsgruber & lars erik zeige. 2022. register: language users’ knowledge of situationalfunctional variation– frame text of the first phase proposal for the crc 1412. register aspects of language in situation (realis) 1(1). 1–57. https://doi.org/10.18452/24901. moore, emma. 2021. the social meaning of syntax. in lauren hall-lew, emma moore & robert j.editors podesva (eds.), social meaning and linguistic variation: theorizing the third wave, 54–79. cambridge: cambridge university press. https://doi.org/10.1017/9781108578684.003. pescuma, valentina n., diana serova, julia lukassek, antje sauermann, roland schäfer, aria adli, felix bildhauer, markus egg, kristina hülk, aine ito, stefanie jannedy, valia kordoni, milena kühnast, silvia kutscher, robert lange, nico lehmann, mingya liu, beate lütke, katja maquate, christine mooshammer, vahid mortezapour, stefan müller, muriel norde, elizabeth pankratz, angela g. patarroyo, ana-maria plesca, camilo rodrı́guez-ronderos, stephanie rotter, uli sauerland, gohar schnelle, britta schulte, gediminas schüppenhauer, bianca maria sell, stephanie solt, megumi terada, dimitra tsiapou, elisabeth verhoeven, melanie weirich, heike wiese, kathy zaruba, lars erik zeige, anke lüdeling & pia knoeferle. 2023. situating language register across the ages, languages, modalities, and cultural aspects: evidence from complementary methods. frontiers in psychology 13. https://doi.org/10.3389/fpsyg.2022.964658. r core team. 2024. r: a language and environment for statistical computing. r foundation for statistical computing vienna, austria. https://www.r-project.org/. rotter, stephanie & mingya liu. 2024. a register approach to negative concord vs. negative polarity items in english. linguistics https://doi.org/10.1515/ling-2023-0016. rotter, stephanie & mingya liu. 2025. less formal and more rebellious — an experiment on the social meaning of negative concord in american english. in chiara gianollo & johan van der auwera (eds.), negative concord: a hundred years on, vol. 385, 303–332. berlin, bosten: de gruyter mouton. rstudio team. 2024. rstudio: integrated development environment for r. rstudio, pbc boston, ma. http://www.rstudio.com/. yalcin, seth. 2007. epistemic modals. mind 116(464). 983–1026. https://doi.org/10.1093/mind/fzm983. zeijstra, hedde. 2007. modal concord. proceedings of semantics and linguistic theory (17). 317–332. https://doi.org/10.3765/salt.v17i0.2961. proceedings of elm 3: 224-235, 2025 mingya liu and stephanie rotter: linguistic and social meaning match: an experiment on modal concord in english. 235 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ is there a superlative rabbit in the ordinal hat? a study of ordinals vs. degree modifiers in nested definites elizabeth coppock, david beaver, wilder perkins, & emily richardson* abstract. this study probes how the semantics of ordinals relates to the semantics of comparatives and superlatives. we examine this question with the help of a picture task in which participants are asked to locate objects described by nested descriptions like the candle on the first/closer/closest table, with an ordinal, comparative or superlative modifier in the inner noun phrase. we show that ordinals systematically lack the ‘relative readings’ observed for unmodified nested descriptions like the rabbit in the hat, in which the inner definite is understood with enriched content, as in the rabbit in the hat with a rabbit in it, in contrast to superlatives. our explanation for this relies on the idea that an ordinal expects an ordering that can be provided by context. keywords. relative readings; nested definite descriptions; ordinals; superlatives 1. introduction. nested descriptions like the rabbit in the hat, have been observed to have ‘relative readings’, in which the inner definite is understood with enriched content, as in the rabbit in the hat with a rabbit in it (haddock 1987). as bumford (2017) observes and explains via scope movement, nested descriptions with superlatives like the rabbit in the biggest hat have relative readings too, in this case paraphrasable as the rabbit in the biggest hat with a rabbit in it. in this paper, we show that ordinals resist relative readings to a substantially greater extent than superlatives (and we find that comparatives easily allow them in our experimental setting). we establish this with the help of a picture task in which participants are asked to locate objects described by nested descriptions like the candle on the first/closer/closest table, with an ordinal, comparative or superlative modifier in the inner noun phrase. the differences we observe are in line with prior work showing differences between ordinals and superlatives (bylinina et al. 2014). however, the results present difficulties for accounts of the semantics of ordinals on which they are entirely parallel to (bhatt 2006) or contain superlatives (alstott 2023). such accounts would predict relative readings with both ordinals and superlatives in nested descriptions, contra what we found in the experiments we will report. likewise, our results show that there is a danger that a model like bumford’s would overgenerate if applied to ordinal descriptions. his type-shifting model, which accounts for the relative reading with superlatives by allowing part of a nested description to be be interpreted with scope above its containing description, must somehow be restricted to prevent nested ordinals from undergoing the same scope-changing operations as nested superlatives. we discuss two general strategies for explaining the contrast we observe between ordinals and superlatives. the first strategy builds on bylinina et al.’s idea that ordinals do not undergo scope movement, and the second builds on the idea that ordinals depend on a contextually salient linear ordering with a basis that is preferably iconic to the natural numbers. *thanks to the elm community for valuable feedback on this work and the opportunity to develop it. authors: elizabeth coppock, boston university (ecoppock@bu.edu), david beaver, university of texas at austin (dib@utexas.edu), wilder perkins, wesleyan university (wperkins@wesleyan.edu) & emily richardson, boston university (erichardson@bu.edu). proceedings of elm 3: 117-129, 2025 c©2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson published by the lsa with permission of the author(s) under a cc by license. 117 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 2. experiments. 2.1. experiment 1. the purpose of experiment 1 (and indeed of both experiments) is to determine whether degree modifiers and ordinal modifiers differ in their propensity to give rise to relative readings for nested definites containing them. ordinals are compared here to two types of degree modifiers: comparatives and superlatives. 2.1.1. materials and design. in both of our experiments, participants were presented with a display involving objects placed on a sequence of locations, along with a series of questions about the display. in experiment 1, participants were presented with the display in figure 3, a set of stairs with various objects placed on them. (in experiment 2, as we will show later, the display featured a sequence of tables rather than stairs.) item n prompt in list {a,b} e1 3 what’s beside the cactus on the {first,lowest} stair? e2 3 what’s beside the cat on the {lowest,third} stair? e3 3 what’s beside the mug on the {third,highest} stair? e4 3 what’s beside the flower on the {highest,first} stair? e5 2 what’s beside the lamp on the {second,higher} stair? e6 2 what’s beside the basketball on the {higher,first} stair? e7 2 what’s beside the apple on the {first,lower} stair? e8 2 what’s beside the fishbowl on the {lower,second} stair? m1 3 what objects are on the {lowest,first} stair with a cactus on it? m2 3 what objects are on the {third,lowest} stair with a cat on it? m3 3 what objects are on the {highest,third} stair with a mug on it? m4 3 what objects are on the {first,highest} stair with a flower on it? m5 2 what objects are on the {higher,second} stair with a lamp on it? m6 2 what objects are on the {first,higher} stair with a basketball on it? m7 2 what objects are on the {lower,first} stair with an apple on it? m8 2 what objects are on the {second,lower} stair with a fishbowl on it? figure 1: exp. 1 display and prompts in lists {a,b}. proceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 118 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ targets. on target trials, the prompt contained a nested definite, i.e., a definite description containing another. the nested definites were of two possible forms, depending on the position of the target adjective (and the noun it modifies): • embedded (the adjective modifies the lower np): what’s beside [dp the cactus on [dp the first stair ] ]? • matrix (the adjective modifies the higher np): what objects are on [dp the first stair with [dp a cactus on it ] ]? the constructions with the adjective in the embedded position are the ones that are potentially revealing as to whether relative readings are available for the adjective. the matrix constructions serve as a control, to ensure that there is nothing conceptually wrong with the meaning that would result, and in particular that there is nothing wrong with the restriction on the comparison class that would be involved in a relative reading. the materials were divided into two lists (a and b), so that no participant saw the same adjective/noun pair in both embedded and matrix positions, but across participants, judgments were collected on both variants. comparative attributives (as in the higher stair with a fishbowl on it) are most felicitous in contexts where there are only two satisfiers of the modified description (stair with a fishbowl on it in this case).1 therefore, with a view toward felicitous use of comparative attributives, the display was set up so that four object types (namely apple, basketball, fishbowl, and lamp) are instantiated exactly twice in the scene. four other object types (cactus, cat, flower, and mug) have three instances in the display, making perfectly felicitous a superlative attributive modifier such as the highest stair with a cat on it. the questions presented to the participants along with the display made reference to one of these eight object types, and cardinality of the object type in the display was coded as the variable cardinality. the prompts varied with respect to modifier type: ordinal vs. degree. ‘degree’ includes both comparative (higher) and superlative (highest) modifiers. whether a degree modifier took the form of a comparative or a superlative was determined by cardinality: • with nouns associated with cardinality 2, degree modifiers were comparative (e.g. what’s beside the lamp on the higher stair?) • with nouns associated with cardinality 3, degree modifiers were superlative (e.g. what’s beside the mug on the highest stair?) modifiers also varied with respect to orientation: • maximal: highest/higher, third/second • minimal: lowest/lower, first the precise choice of ordinal modifier in this experiment depends on cardinality, too: the maximal ordinal for cardinality 2 is second; for cardinality 3 it is third. 1aparicio et al. (2021) show that attributive comparative modifiers can be felicitous in contexts with more than two satisfiers of the modified description. however, their findings still support the notion that comparison classes must be binary in some sense, as the bigger circle was only felicitous when there were only two sizes of circles represented in the display. they propose that the binarity requirement that comparatives impose is really on the number of degrees instantiated in the context rather than on the number of objects. here that distinction makes no difference. proceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 119 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the experiment thus had a fully crossed 2×2×2×2 design, for every combination of the two levels of position, modifier type, cardinality, and orientation, although orientation was not a treatment variable of interest, and was only included for variety among the stimuli. the cases with comparatives are those where modifier type is degree and cardinality is 2; the cases with superlatives are those where modifier type is degree and cardinality is 3. there were two lists of target stimuli, both containing 16 questions, 8 with the modifier in matrix position, and 8 with the modifier in embedded position. each set of 8 covered all four cardinality-2 nouns and all four cardinality-3 nouns. in each such group of four, two involved ordinal modifiers (one maximal, one minimal) and two involved degree modifiers (one maximal, one minimal).2 no noun appeared with the same modifier twice in the same list. the two lists are spelled out in the table in figure 3. the target and filler lists were independently shuffled randomly, and then interleaved with each other evenly so that every target trial was preceded by a filler trial. the objects were placed carefully in the display so as to avoid absolute readings of the embedded modifiers in the stimuli. for prompts with lower or lowest, this means that the noun’s object type cannot be instantiated on the bottom stair; mutatis mutandis for higher and highest and the top stair. with ordinal prompts, we intend the lowest stair to be the “first” stair, but there is a risk that the participant counts down from the top instead, treating the topmost stair as the “first”, and the one below that as the “second”. to guard against absolute readings arising through this direction, stimuli involving first never invoked an object type that was instantiated on the top stair; mutatis mutandis for second and third. in addition, we ensured that the each object had a distinctive set of shelfmates. this enabled the free response texts provided by the participants to be coded as to whether the participant had identified the target referent. familiarization stimuli. for the purpose of familiarization with the display prior to answering the target questions, participants were asked to study the display, and answer questions like how many mugs are there? for the nine object types appearing in the display: mugs, apples, flowers, candles, cats, fish bowls, cactuses, lamps, and basketballs. the familiarization phase also included some what color questions: • what color is the ball beside the lowest cactus? • what color is the mug beside the third cactus? • what color is the cat beside the higher fishbowl? • what color is the apple beside the first cat? • what color is the candle beside the cat? • what color is the mug beside the fishbowl? the last two carry false presuppositions (there is no candle beside the cat, or mug beside the fishbowl), so the participants were expected to write “doesn’t make sense” for these cases. 2partially in response to the need to avoid absolute readings, it turned out that the nested definite in the a list was usually but not always coreferential with the corresponding nested definite in the b list. for example, with item e8, the a list has lowest (minimal) and the b version has third (maximal). however, this did not lead to a skew in the materials with respect to orientation. in both lists, there were two items for each position × type × cardinality condition, one maximal and one minimal. proceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 120 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ item prompt f0 what else is on the stair that has a mug and a lamp? f1 what else is on the stair that has a cat and a basketball? f2 *what else is on the stair that has a basketball and an apple? f3 what else is on the stair that has a cactus and a fishbowl? f4 *what object is between a flower and a candle? f5 what object is between a basketball and a fishbowl? f6 what object is between a lamp and a cactus? f7 what object is between a cat and a mug? f8 what object is between an apple and a mug? f9 how many stairs have both an apple and a lamp? f10 how many stairs have both a flower and a fishbowl? f11 how many stairs have both a flower and a cactus? f12 how many stairs have both a basketball and a candle? f13 how many stairs have both a cat and a mug? f14 *what else is on the stair that has a lamp and a fishbowl? f15 what else is on the stair that has a mug and an apple? f16 what else is on the stair that has a flower and a cactus? table 1: filler stimuli for experiment 1. fillers. fillers (listed in table 1) were included among the target stimuli in order to introduce variety and provide some clear cases where “doesn’t make sense” could be used, so that participants would solidify their confidence in deploying that option. examples marked with an asterisk in table 1 are ones to which a “doesn’t make sense” response is expected. 2.1.2. outcome variable coding. the free text responses provided by a participant would generally be in the form of a list of object types. the objects listed allowed us to identify which referent the participant assigned to the nested definite, and in particular whether the participant had assigned a relative reading to the adjective. we used a three-way coding of the responses: • 1 rejection of relative reading • 0 acceptance of relative reading • na unexpected response (acceptance of an unexpected reading) we coded the response as 1 (rejection) if the participant wrote some variant on “doesn’t make sense”. we coded the response as 0 (acceptance) if the participant listed the items that accompany the referent on a relative reading of the adjective. examples of responses that were coded as na include the following: • prompt: what objects are on the third stair with a cat on it? response: apple, flower, mug (these are on the second stair, which has a cat on it) expected under relative reading: fishbowl proceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 121 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ • prompt: what objects are on the first stair with a basketball on it? response: a grey cat (maybe they thought the fishbowl was a basketball and counted from the top?) expected under relative reading: flower, cactus, fishbowl in these cases, participants identify referents that are incompatible with the descriptions’ literal meanings. 2.1.3. participants. 40 native speakers of english were recruited via prolific. 2.1.4. results. we coded 22 of the 640 responses to target questions as na (3.2%). for the purposes of the statistical analysis, the na responses were removed, leaving a binary outcome variable, 1 for ‘reject’ and 0 for ‘accept’. the mean rejection rates in each of the eight conditions of interest in experiment 1 are shown in figure 2. the group means and associated confidence intervals are listed in table 2. with the modifier in matrix position (first/lowest stair with a cat), there was almost no rejection (as expected). degree modifiers in the matrix position were rejected under 4% of the time, and indeed the rejection rate for these conditions was not significantly different from zero, as shown by the fact that the 95% confidence intervals cross the zero line. with ordinals in the matrix position, rejection was rare (under 25%), although the rejection rate is not entirely negligible, and is significantly different from zero, as shown by the fact that the confidence intervals around the means for these conditions do not cross the zero line. despite this wrinkle, the matrix examples did establish a low-rejection baseline relative to which the embedded examples can be evaluated. the effect of ordinal vs. degree modifier is strikingly large in the embedded condition, where relative readings reveal themselves. a strong majority of respondents (over 80%, in both cardinality conditions) rejected relative readings for nested descriptions with ordinal modifiers in the embedded position (cat on the first stair). relative readings for nested descriptions containing degree modifiers were sometimes rejected, but less often than with ordinals. moreover, rejection of a relative reading in the embedded position was over 30 percentage points more common with superlatives (45.58%) than with comparatives (10.25%). this contrast is reflected by the difference in height of the red bars in the upper two quadrants of the graph in figure 2. the rejection rate for relative readings of superlatives in embedded position is shown by the red bar under ‘embedded 3’; the corresponding rate for comparatives is shown by the red bar under ‘embedded 2’. the results of statistical tests confirm the visual impressions that figure 2 gives rise to. using the lmer and lmer-test packages in r, we constructed a mixed-effects logistic regression model with rejection as the outcome variable and position, modifier type, and cardinality along with all possible twoand three-way interactions as fixed effects and participant as a random effect. an anova test based on that model revealed a significant interaction of position×modifier type (p < 0.001), showing that the contrast between ordinals and degree modifiers is reliably greater in the embedded position. there was also a significant 3-way interaction between position, modifier type, and cardinality (p < 0.05), reflecting the fact that relative readings of superlatives in the embedded position were rejected more often than relative readings of comparatives in that position. the model coefficients and associated p-values are given proceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 122 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 2: results of experiment 1. error bars show 95% ci. in table 3. the orientation (lower vs. higher, first vs. second, etc.) of the modifier was not a treatment variable of interest in this experiment, but given that it was systematically manipulated, we had the opportunity to see whether rejection was driven primarily by modifiers of one orientation or another. given that first has been argued to be a superlative underlyingly (alstott 2023), we might expect that the maximal ordinals second and third would be driving the high rejection rate for ordinals in the embedded position. we therefore built a mixed-effects logistic regression model on the embedded position subset of the data with modifier type, cardinality, and orientation along with all possible two and three-way interactions, and used anova tests to determine which of these factors has a significant effect on rejection rate. although there was a trend whereby maximal ordinal modifiers were slightly more likely to induce rejection than minimal ones, no significant effect of orientation was detected. 2.1.5. discussion. overall, relative readings were found to be far less accessible with ordinal modifiers than with degree modifiers. the contrast is particularly stark when there are two objects of the relevant type: the lamp on the higher stair has a far greater chance of referring successfully proceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 123 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ position modifier type # estimate std. error ( conf.low conf.high ) embedded degree 2 0.1025 0.0345 0.0337 0.1714 embedded degree 3 0.4358 0.0565 0.3233 0.5484 embedded ordinal 2 0.8076 0.0449 0.7182 0.8971 embedded ordinal 3 0.8461 0.0411 0.7642 0.9280 matrix degree 2 0.0128 0.0128 -0.0127 0.0383 matrix degree 3 0.0375 0.0213 -0.0050 0.0800 matrix ordinal 2 0.2027 0.0470 0.1089 0.2964 matrix ordinal 3 0.2162 0.0481 0.1201 0.3120 table 2: group means (average rates of rejection) for each condition in experiment 1 estimate sum sq mean sq df f value pr(>f) sig. intercept 0.0405 card=3 -0.0375 0.0312 0.0312 1/592 0.2785 5.98e-01 type=ordinal 0.0879 11.9345 11.9345 1/592 106.5541 4.40e-23 *** pos=emb 0.2500 49.4699 49.4699 1/592 441.6781 1.11e-73 *** card=3×type=ord 0.0501 1.4422 1.4422 1/592 12.8765 3.60e-04 *** card=3×pos=emb 0.2933 0.1111 0.1111 1/592 0.9923 3.20e-01 type=ord×pos=emb 0.5620 4.1152 4.1152 1/592 36.7422 2.40e-09 *** card=3×type=ord×pos=emb -0.4809 2.3024 2.3024 1/592 20.5571 7.01e-06 *** table 3: mixed effects logistic regression model coefficients and associated p-values based on anova tests for experiment 1. signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1. to the lamp on the fourth stair than the lamp on the second stair does. this effect was not driven solely by ordinal modifiers of maximal orientation (like second, when there are two objects of the relevant type); absolute readings for first were rejected at comparable rates. the fact that rejection was significantly more common with superlatives than with comparatives may be related to the availability of absolute readings for the inner dp. in the case of superlatives, the inner dp can be used in isolation to refer to one of the stairs, even though such a parse leads ultimately to global reference failure. for example, the lowest stair is a felicitous way of referring to the bottom stair. in contrast, in a context with many stairs, the lower stair does not refer in isolation, because comparatives require binary comparison classes, as mentioned above. participants who reject nested descriptions with superlatives in the embedded position may be fixating on an absolute reading. aparicio et al. (under revision) refer to this manner of being led astray by a doomed assignment of a referent as a ‘referential garden path’ effect. with comparatives, there is no absolute reading to fixate on in this context, loosening the grasp of this reading on the listener’s mind, thereby making the relative reading more available. it would not be entirely straightforward to test this hypothesis. the only type of context in which an absolute reading would be available for comparatives would be one in which the total number of stairs is 2, and then we would not be able to test for relative readings (and superlatives would be less optimal). eliminating the absolute reading for superlatives would not be possible proceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 124 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ as long as there is a finite set of stairs (or equivalent) with a lowest and a highest. perhaps if the display were a snapshot of a staircase that continued below and above the visible region, then an absolute reading for a superlative could be eliminated. in that case, we would expect superlatives to be on par with comparatives in the embedded condition, with a negligible rejection rate, reflecting the availablity of relative readings. we leave this as a prediction to test in future research. 3. experiment 2. to estimate the robustness of the effect found in experiment 1, and to see whether the effect generalizes from vertical sequences like stairs to horizontal ones, we carried out a variation on that experiment using a horizontal sequence of tables, as shown in figure 3. 3.1. materials and design. like experiment 1, this experiment compares the availability of relative readings of nested definites with ordinal attributive modifiers (as in the second table with a fishbowl on it) with degree attributive modifiers, including both comparative and superlative attributives. the display and the prompts were designed to ensure that an absolute reading of the nested definite would not be available, leaving only a relative reading. rejection (“doesn’t make sense”) in this environment thus signalled the absence of a relative reading. the treatment variables were the same as in experiment 1. along with examples with modifiers in the embedded position, corresponding examples with the same modifiers in matrix position were included as controls (modifier position: matrix vs. embedded). as in the previous experiment, the object types picked out by nouns in the prompts varied according to their cardinality, being instantiated either twice or three times. comparatives were used with cardinality-2 nouns and superlatives were used with cardinality-3 nouns. here, the degree adjectives used were closer and closest rather than lower and lowest, and farther and farthest instead of higher and highest. the ordinals were the same (first, second, third). as before, the modifiers also varied with respect to orientation: maximal vs. minimal. 3.2. participants. another 40 native speakers of english were recruited via prolific. 3.3. results. there were two responses coded as na (signifying that the participant neither rejected the question nor listed the expected set of shelfmates). these were removed from the dataset for the purposes of the analysis, leaving us with a binary response variable, 1 for rejection and 0 for acceptance. the rejection rates in each of the conditions of interest in in experiment 2 are shown in figure 4 (righthand panel), which also repeats the results from experiment 1. as the graph shows, we found the same pattern of results as in experiment 1. with the modifier in matrix position (first table with a cat), there was almost no rejection. in fact, with superlatives in matrix position, rejection never occurred. a strong majority of respondents rejected relative readings for nested descriptions with ordinal modifiers in the embedded position (cat on the first table). relative readings for nested descriptions containing degree modifiers were sometimes rejected, but less often than with ordinals. and again, rejection was more common with superlatives than with comparatives. statistical analysis confirmed that these contrasts are reliable. as before, we constructed a mixed-effects logistic regression model with participant as a random effect and fixed effects for modifier position, modifier type, and cardinality, along with all possible twoand three-way interactions of those factors. anova tests based on this model revealed significant efproceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 125 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ item # prompt for {a,b} list e1 3 what’s next to the cactus on the {first,closest} table? e2 3 what’s next to the cat on the {closest,first} table? e3 3 what’s next to the mug on the {third,farthest} table? e4 3 what’s next to the flower on the {farthest,third} table? e5 2 what’s next to the lamp on the {second,farthest} table? e6 2 what’s next to the candle on the {farther,second} table? e7 2 what’s next to the fish bowl on the {first,closer} table? e8 2 what’s next to the apple on the {closer,first} table? m1 3 what’s on the {closest,first} table with a cactus on it? m2 3 what’s on the {first,closest} table with a cat on it? m3 3 what’s on the {farthest,third} table with a mug on it? m4 3 what’s on the {third,farthest} table with a flower on it? m5 2 what’s on the {farthest,second} table with a lamp on it? m6 2 what’s on the {second,farther} table with a candle on it? m7 2 what’s on the {closest,first} table with a fish bowl on it? m8 2 what’s on the {first,closer} table with an apple on it? figure 3: exp. 2 display and prompts fects of modifier position and modifier type individually, an interaction between modifier position and modifier type, an interaction between modifier type and cardinality, and a three-way interaction among these three (all at the the p < 0.001 level). 4. discussion. we conclude that ordinals are substantially less susceptible to relative readings than degree modifiers, in nested descriptions. this effect was found to be robust across modifier orientations (picking out the lowest/closest/first object vs. the highest/farthest/second/third object), and across orientations of the display (horizontal vs. vertical). granted, the comparison between ordinals and comparatives is not entirely fair because with ordinals there is a competing absolute reading for the inner definite, while with comparatives there is not. if the contrast between comparatives and superlatives is due to the availability of an absolute reading for the inner definites, then the difference in rejection rates for those two types of modifiers can perhaps be viewed as a measure of the magnitude of that effect. the contrast between superlatives and ordinals that goes over and above that cannot be attributed to the availability of an absolute reading; it shows that there is something else preventing relative readings for ordinals. the real finding of interest from this paper is the contrast between superlatives and ordinals, because these are fairly matched in terms of the existence of the absolute reading. one strategy for explaining this result is to adopt bylinina et al.’s stipulation that ordinals cannot undergo scope movement, made in order to explain the absence of ‘upstairs de dicto’ readings proceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 126 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 4: results of experiments 1 (left) and 2 (right). error bars show 95% ci. position mod.type # estimate std.error (conf.low conf.high) embedded degree 2 0.2875 0.0509 0.1861 0.3888 embedded degree 3 0.5443 0.0563 0.432 0.6565 embedded ordinal 2 0.9375 0.0272 0.8832 0.9917 embedded ordinal 3 0.7625 0.0478 0.6672 0.8577 matrix degree 2 0.0375 0.0213 -0.0051 0.0801 matrix degree 3 0 0 0 0 matrix ordinal 2 0.1265 0.0376 0.0516 0.2015 matrix ordinal 3 0.1392 0.0391 0.0612 0.2172 table 4: group means for each condition in experiment 2. with ordinals. this assumption alone does not suffice to block relative readings, though, because in order to generate focus-related relative readings of ordinals as in bhatt’s (2006) johnf gave the first telescope to mary, bylinina et al. assume that ordinals expect an implicit comparison class. so under this approach, one would need a theory of why the comparison class argument of first in the cat on the first table cannot be set to ‘with a cat on it’. another possibility is related to constraints on the comparison class that regulate the choice between interpretations. on this view, ordinals and superlatives are subject to both absolute and relative readings. we have seen that even though superlatives can have relative readings, such readings are sometimes rejected when an absolute reading is available in the context. this suggests that there can be a stubborn adherence to the absolute reading, so stubborn that it will accept reference failure, even though a relative reading is generated by the grammar. the idea that we would like to put forth is that given a choice between comparison classes, listeners will strongly prefer a comparison class that has an iconic basis, and in particular a basis that is iconic to the natural numbers. suppose that an ordinal expects an ordering that can be proceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 127 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ provided by context. the ordering is a function f from a ‘basis’ to satisfiers of the modified predicate. the basis is a linearly ordered set like a sequence of times (as in second train) or locations (second stair). the nth table is the nth object in a sequence ⟨f(i1), f(i2), f(i3), . . .⟩. we posit further that the more iconic a sequence is to the natural numbers, the more accessible it is as a basis for the ordering. the more evenly spread out a sequence is, as measured by a perceptually salient distance metric, the more iconic it is to the natural numbers. a visual display that is iconic to the natural numbers, such as a series of stairs or tables, will satisfy that iconicity constraint. the comparison classes that one is forced to construct in the matrix constructions in our experiments, on the other hand, restricting to only those stairs with a cat on them, do not have a basis that is iconic to the natural numbers. the low rejection rates in the matrix conditions show that bases need not be iconic, but in those cases the listener has no choice about how to construe the basis. in the embedded condition, the listener’s task is to choose a scope interpretation for the modifier. the relative reading depends on a non-iconic basis and the absolute reading depends on an iconic one. we suggest that in this experiment, the iconicity of the visual display pulls so strongly in favor of the absolute reading for the ordinals that the relative reading cannot be accessed. on this view, scope movement is possible in principle for ordinals, and is predicted to manifest itself in circumstances where competition from an absolute reading is not as great. in closing, our consideration of relative readings for ordinals and degree modifiers shows that the pattern of availability of relative readings in haddock descriptions is quite intricate. there are many types of operator for which there has, as yet, been no systematic study of their behavior in these constructions. these include exclusives: our judgment is that the rabbit in the only hat and the mug on the only stair lack relative readings, and would be infelicitous in cases where there are multiple salient hats or stairs, respectively. the behavior of exclusives thus appears to parallel that of ordinals, so that once again a scope-based explanation of relative readings is in danger of overgenerating. other expressions to consider in haddock descriptions are markers of identity like other in the mug on the other step, further types of scale-dependent expressions (e.g. next in the mug on the next step), and proportional quantifiers, like every and most in the mug on every/most step(s). note that with regard to quantifiers, the standard wide scope universal reading is distinct from the relative reading the mug on every/most step(s) (with a mug on it/them), and could not be generated by the same mechanism bumford (2017) uses for definite descriptions. it is clear that haddock descriptions present an empirically rich and puzzling domain for which further empirical investigation can potentially shed light on a range of interconnected theoretical issues, not only ordinal and degree operator semantics, but also the nature of scope-taking, and contextual restriction more broadly. references alstott, johanna. 2023. ordinal numbers: not superlatives, but modifiers of superlatives. paper presented at salt 33. aparicio, helena, curtis chen, roger levy & elizabeth coppock. 2021. granularity in the semantics of comparison. in nicole dreier, chloe kwon, thomas darnell & john starr (eds.), proceedings of semantics and linguistic theory (salt 31), 550–569. 10.3765/salt.v31i0.5121. aparicio, helena, roger levy & elizabeth coppock. under revision. referential garden path efproceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 128 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ fects in modified haddock descriptions. ms., cornell university, mit, and boston university. bhatt, rajesh. 2006. covert modality in non-finite contexts. berlin: mouton de gruyter. bumford, dylan. 2017. split-scope definites: relative superlatives and haddock descriptions. linguistics and philosophy 40(6). 549–593. bylinina, lisa, natalia ivlieva, alexander podobryaev & yasutada sudo. 2014. a non-superlative semantics for ordinals and the syntax and semantics of comparison classes. haddock, nicholas j. 1987. incremental interpretation and combinatory categorial grammar. in proceedings of the 10 international joint conference on artificial intelligence, vol. 2, 661– 663. morgan kaufmann. proceedings of elm 3: 117-129, 2025 elizabeth coppock, david beaver, wilder perkins, and emily richardson: a study of ordinals vs. degree modifiers in nested definites. 129 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ constructing focus alternatives from context and the limits of semantic priming christian muxica & jesse a. harris* abstract. interpreting focus requires a comprehender to identify the set of alternatives intended by the speaker. previous psycholinguistic research has characterized this process in terms of a two-stage model that initially forms an alternative set via the context-insensitive mechanism of semantic priming (gotzner et al. 2016, husband & ferreira 2016). we have instead advanced a one-stage immediate-access model, in which alternatives are immediately constructed from the discourse context (muxica & harris to appear). in two cross-modal probe recognition task experiments, we further test our prediction that the discourse context strongly influences response speed at early moments of focus interpretation. the results are interpreted as uniquely supporting the immediate-access model. keywords. focus; alternatives; sentence processing; probe recognition 1. introduction. a long tradition of research in semantics addresses the interpretive effect of the focus of an utterance. in general, focus serves to highlight an element in some way against the background of the discourse (krifka 1992). focus typically refers to a constituent that is marked by prosodic, syntactic, or, in some languages, morphological means. as focus can be ambiguous in text, we mark it here with the f-feature, as in (2) below. since at least jackendoff (1972), focus has been understood as evoking a contextually salient set of alternative expressions. rooth’s (1985, 1992) alternative semantics framework formalizes the alternative set as the focus semantic value, consisting of expressions of the same formal semantic type that can be substituted for the expression in focus. while focus itself does not alter the truth value of an utterance directly, focus can nonetheless affect the inferences associated with the utterance and, in the case of focus-sensitive semantic operators like only and even, determine the background against which those operators are interpreted (e.g. beaver & clark 2009). multiple kinds of focus have been identified in the literature, including informational and contrastive, among others (for overview, see büring 2016). informational focus marks non-given constituents of an utterance, typically, though not exclusively, as new information. such uses can be illustrated with question-answer pairs, as in (1). in b’s reply, the element in focus (willie) provides new information while the remainder of the utterance is given in the discourse, in that it is information that has been previously mentioned. (1) a. who did dolly sing to? b. dolly (only) sang to [willie]f in (1), the set of relevant alternatives to the focused word willie depends on the context. it might be limited to members of dolly’s band, other country stars, a particular concert, or even humanity *acknowledgments: we would like to thank audiences at the third meeting of experiments in linguistic meaning and the ucla psycholinguistics/computational linguistics seminar. we would also like to thank our wonderful team of undergraduate ras at the ucla language processing lab. last, but certainly not least, we are indebted to thomas bye for generously funding this and other research projects. authors: christian muxica, university of california los angeles (cmuxica@g.ucla.edu) & jesse a. harris, university of california los angeles (jharris@humnet.ucla.edu). proceedings of elm 3: 262-275, 2025 c©2025 christian muxica and jesse a. harris published by the lsa with permission of the author(s) under a cc by license. 262 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ at large. in the absence of a focus operator like only, b’s answer suggests that dolly did not sing to anyone else relevant in the set of alternatives, though the strength and cancelability of such exhaustive implicatures is perhaps debatable. on the other hand, contrastive focus marks constituents that stand in polar contrast with an item that either has been mentioned previously in discourse or is highly accessible. in (2), for instance, the focus on willie evokes merle as an alternative. consequently, speaker b’s utterance not only expresses the entailment that dolly sang to willie, but again the exhaustive implicature that dolly did not sing to merle. this implicature arises despite the fact that speaker b did not explicitly mention merle. instead, the preceding discourse context makes merle a highly salient alternative. in the remainder of this paper, we concentrate on contrastive focus, with the expectation that many of our conclusions would extend to other types of focus, as well. (2) a. i heard that dolly sang to merle b. dolly (only) sang to [willie]f successful interpretation of any utterance containing focus requires a comprehender to identify the set of alternatives intended by the speaker. as mentioned, the relevant set of alternatives is determined by the context, and, as yet, it is unclear exactly how the set of alternatives might be identified by comprehenders. in recent years, a novel body of psycholinguistic research has characterized this process as a two-stage model that crucially relies on lexical activation to generate and select alternatives (gotzner et al. 2016, husband & ferreira 2016). in the first, contextinsensitive, stage, lexical associates – i.e., words which bear a strong lexical association with the word in focus, become highly activated through semantic priming immediately upon encountering a focused constituent. in the second stage, a context-sensitive mechanism identifies the relevant alternatives from among these associates and maintains their activation. eventually, the activation of non-alternatives will decay, leaving the relevant alternative set behind.1 as we understand it, the two-stage (sometimes known as the alternative activation) model advances two main hypotheses about the selection of alternatives, which we refer to as (i) priming dependence and (ii) late generation. priming dependence is the hypothesis that constructing a representation of the alternative set depends on semantic priming from the element in focus. late generation is the hypothesis that additional time, after the focus is encountered, is required to select only those alternatives that are contextually relevant. as a result, the two-stage model can be considered a destructive model of the alternative set, in which alternatives are first proposed by the lexicon and then disposed of by context. however, we have previously argued against both of these claims. in muxica & harris (to appear), we proposed the immediate-access model, in which membership in the alternative set is guided by contextual constraints immediately after focus is encountered, without first being mediated by semantic priming. therefore, priming and the generation of focus alternatives can be separated conceptually, as alternatives are constructed immediately in concert with the context. for this study, we adapted the materials and design from muxica & harris (to appear) in order to further probe the priming independent aspect of the immediate-access model. we tested 1it should be noted that husband & ferreira (2016) remain agnostic as to whether the second stage involves a passive process of decay or an active process of suppression. for ease of presentation, we will exclusively describe the two-stage model in terms of decaying activation. proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 263 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the prediction that response times in a probe recognition task are immediately influenced by the alternative status of a probe word, even when that probe word is not a semantic associate of the word in focus. we explicitly manipulated the alternative status of such non-associate probe words through the discourse context. in doing so, we were able to more directly address the relative contributions of context and semantic priming in the selection of focus alternatives. 2. the two-stage model. by and large, the studies which have investigated the selection of alternatives have supported the properties of priming dependence and late generation. multiple cross-modal forced-choice task experiments have yielded results compatible with the idea that contextually relevant alternatives are activated alongside associate non-alternatives in the earliest moments of processing focus (husband & ferreira 2016, gotzner et al. 2016, gotzner & spalek 2019, lacina et al. 2023, jördens et al. 2020, braun & tagliapietra 2010). until recently, only with additional time has an advantage for contextually relevant alternatives alone been observed (husband & ferreira 2016, gotzner et al. 2016). three foundational studies from this literature are reviewed below. husband & ferreira (2016) conducted a cross-modal lexical decision experiment with a betweensubjects stimulus onset asynchrony (soa) manipulation. on each trial, subjects listened to an utterance and then responded to one of three lexical decision targets (3-b). for half of the subjects, targets appeared immediately after (i.e., at 0ms soa) the focused word (sculptor). for the other half, targets appeared after a brief 750ms delay. target words differed in their relationship to the focused word. targets were either plausible associate alternatives (painter), implausible associate alternatives (statue), or implausible non-associate alternatives (register), included as a control. (3) a. sample item from husband & ferreira (2016) the museum thrilled the [sculptor]f . . . b. lexical decision targets alternative: painter associate non-alternative: statue control: register at probe points presented immediately after the focused word, husband & ferreira (2016) found a simple effect of semantic priming: subjects responded faster to associates of the focus (painter and statue) than non-associates (register). however, after a 750ms delay, there was an effect of focus: subjects responded faster to the associate alternative (painter) than either the nonalternative (statue) or the control (register). these results are clearly compatible with the two-stage model. the earliest moments after encountering focus appear to be context-insensitive purely reflecting the semantic priming of the first stage. after a delay, the activation of irrelevant items decays, revealing the context-sensitive selection of the second stage. one limitation of husband & ferreira (2016) is that their materials do not include an explicit discourse context. as discussed, the alternative status of any given element is primarily determined with respect to contextual relevance. however, the alternative and non-alternative status of targets in husband & ferreira (2016) was determined purely with respect to plausibility, which is only one component of contextual relevance. the experiments in gotzner et al. (2016) and gotzner & spalek (2019) addressed this conproceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 264 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ cern in two cross-modal probe recognition experiments in german with a similar between-subjects soa manipulation. their materials consisted of two speaker dialogues. the first speaker introduced a set of alternatives (peaches, cherries, and bananas) relevant for the focus used by the second speaker (peaches). on each trial, subjects listened to these dialogues, and then were presented with one of three probe words, indicating whether or not they heard that probe word in the preceding dialogue. in gotzner & spalek (2019), the probe recognition task was administered immediately after (i.e., 0ms soa) the focused word (peaches). in gotzner et al. (2016), the task was administered following a 2050ms delay. the probe words again differed in their relationship to the focused word. probe words were either associate mentioned alternatives (cherries), associate unmentioned alternatives (melons), or non-associate unmentioned non-alternatives (clubs), included as a control. (4) a. sample dialogue from gotzner et al. (2016) a. in the fruit bowl, there are peaches, cherries, and bananas i bet carsten has eaten cherries and bananas b. no, he only ate [peaches]f b. probe words mentioned: cherries unmentioned: melons control: clubs when tested immediately after the sentence, there was a simple semantic priming effect; subjects responded faster to the associates (cherries and melons) than the non-associates (clubs). after a delay, there was an effect of context such that responses to mentioned alternatives (cherries) were faster than either the unmentioned (melons) or the control (clubs) probe word. these results receive a natural explanation under the two-stage model: effects at the early soa reflect the first stage of context-insensitive semantic priming, while effects at the late soa reflect the context-sensitive selection of focus alternatives in the second stage. additional cross-modal forced choice task studies have investigated the selection of alternatives (e.g., lacina et al. 2023, jördens et al. 2020). for the most part, these studies have argued in favor of the two-stage model. in fact, until recently, the two-stage model was perhaps the only existing model for the selection of alternatives. we now briefly present our recent study which presented evidence against the two-stage model and advances the immediate-access model as an alternative. 3. the immediate-access model. in muxica & harris (to appear), we identified a number of conceptual challenges for the two-stage model. first, focus is an extremely flexible phenomenon; almost any element of the same type-theoretic category as the focus can serve as an alternative given proper contextual support. in cases of broad focus, complex constituents that contain multiple words can be marked for focus, as can entire utterances. at this point, it is not clear how lexical-level priming is meant to generate alternatives in cases of broad focus. second, recall that the two-stage model is priming dependent. this means that the contextsensitive selection of alternatives in the second stage depends upon the lexical activation generated by semantic priming in the first stage. for example, imagine a context in which a group of artist proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 265 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ has painted a mural that depicts a tank driving through a meadow. in such a context, tank is clearly a relevant alternative to flowers in (5). presumably, flowers and tank are semantically unrelated, and thus these words cannot be associates in the two-stage model. (5) a. what did simon paint on the mural? b. simon only painted [the flowers]f on the mural it’s unclear how non-associate alternatives such as tank could enter the alternative set without first entering as a semantic prime under the two-stage model. on these grounds, we have argued that context immediately guides the selection of alternatives in focus processing and that the priming effects observed in past studies are at least partially independent effects of low-level lexical activation, unrelated to the interpretation of focus. a number of implementational questions arise once the priming dependent aspect of the twostage model is relaxed. in muxica & harris (to appear), we addressed the issue of when nonassociate, i.e., semantically unrelated, words are given as alternatives in the context. thirty items consisting of two speaker dialogues were presented to listeners in a cross-modal probe recognition experiment. the first speaker in the dialogue introduced an associate alternative (muffin) and a nonassociate alternative (pistol) relevant for the second speaker’s use of focus (violin). additionally, the first speaker mentioned a non-associate (house) which served as a control. unlike previous studies, all the probe words were given in the discourse, arguably allowing us to better disentangle the effects of focus from the effects of a discourse new word (see also hoeks et al. 2023). these words served as probes in the recognition task and were controlled for various lexical factors (e.g., frequency, number of morpheme, orthographic neighborhood size, etc.). on each trial, subjects performed a probe recognition task immediately (i.e., 0ms soa) after the presentation of focus. (6) sample dialogue from muxica & harris (to appear) a. jonah brought the guitar and the pizza to band practice at the new house b. no, he only brought the [violin]f (7) probe words associate: muffin non-associate: pistol control: movie as expected, responses to probe words were faster for alternatives (muffin and pistol) than for non-alternatives (house). crucially, there was no evidence of a difference between the associate and non-associate alternatives. in fact, bayes factors provided evidence against the hypothesis that response times to associate and non-associate conditions were different. we took these results to be incompatible with the two-stage model and in support of the immediate-access model. we suggested that the previous literature may have had confounded semantic association with alternative status, obscuring the early effect of contextual relevance on focus computation. the immediate-access model crucially predicts that the discourse context is recruited to select alternatives immediately upon encountering focus. however, we did not explicitly manipulate the discourse context to explore another crucial prediction of the model, namely, that focus alternatives are directly determined by contextual information. the current study aims to address this proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 266 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ key prediction. we adapted the design and materials from muxica & harris (to appear), explicitly manipulating the alternative status of the non-associate probe word in the discourse context. to preview, we found that response times to non-associate probe words was modulated by their contextual relevance as alternatives. we take this finding as further evidence that a contextuallyrelevant alternative set is generated immediately after focus is encountered, partially replicating our previous results and further supporting the main predictions of the immediate-access model. 4. experiment. 4.1. materials and method. the 3 condition design from muxica & harris (to appear) was adapted into a 2×2 factorial design (context × probe word). our materials consisted of 28 pairs of audio dialogues and 28 pairs of written probe words. (8) one-alternative (one-alt) context a1. after eating leftover pizza, jonah brought the guitar to band practice at the new house two-alternative (two-alt) context a2. jonah brought the guitar and the pizza to band practice at the new house (9) target sentence for both contexts b. no, he only brought the [violin]f (10) probe words associate: guitar non-associate: pizza written probe words were presented in one of two conditions. in the associate condition, the probe word (guitar) was closely related to the word in focus (violin). while in the non-associate condition, the probe word (pizza) was not closely related to the word in focus. in both conditions, the probe word was mentioned in the preceding audio dialogue. and thus, the correct response to the probe recognition task was always “yes” on critical trials, precluding the possibility of a response bias confound between conditions. filler items balanced the overall distribution of “yes” and “no” responses. audio dialogues consisted of a context sentence in one of two conditions (8), followed by a target sentence (9). in both contexts, speaker a’s utterance described a situation using the associate and non-associate probe words. speaker b responded using corrective associated focus (i.e., no + only). in the response, the focused word provides the only new information as the rest of the content words were previously given in the context. in the two-alt context, the associate and non-associate probe words were conjoined arguments of a main verb (e.g., jonah brought the guitar and the pizza), making both probe words contextually relevant focus alternatives to the word in contrastive/corrective focus (violin). however, in the one-alt context, only the associate probe word (guitar) appeared as an argument of a main verb, making it the only contextually relevant focus alternative in the target sentence. crucially, although the non-associate probe word was not a contextually relevant alternative, it was still mentioned by speaker a in the one-alt condition. the same recordings from the previous study were used for all of speaker b’s utterances and for all of speaker a’s utterances in the two-alt context. as the one-alt context is novel to this proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 267 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ study, twenty-eight new contexts were recorded from speaker a. as in our previous study, speaker b was a male speaker trained in the production and transcription of english intonation using the tones and breaks indices system (tobi; pierrehumbert & hirschberg 1990). he was instructed to use an l+h* pitch accent when producing the final focused word.2 speaker a was the same female speaker from the previous study. she had not received any formal prosodic training previously and was instructed to produce the items naturally, rather than with any specific contour. the 28 pairs of probe words (associate and non-associate) were identical to the previous study and were controlled for length, frequency, number of morphemes, and orthographic neighborhood size (balota et al. 2007, brysbaert & new 2009).3 we also controlled for the semantic association between each probe word and their corresponding focus using latent semantic analysis (lsa; landauer & dumais 1997) and an internet norming study in which ucla undergraduates provided semantic similarity judgements on a 7-point likert scale.4 the same fifty-six two-speaker filler dialogues from the previous study were used, resulting in a final list of 84 items. the probe word was not mentioned in 42 of these filler items in order to balance the distribution of responses across the study. both of our speakers were instructed to produce the filler items naturally rather than with any specific contour. 4.2. analysis. before discussing the results of the pilot and the main in-person studies, the general procedure for data cleaning and analysis is described. all subjects included in the final data set answered at least 75% of questions on the probe task and comprehension questions correctly. only correct responses to the probe recognition task were retained for the response time analysis. response times below 200ms and above 2,500ms were excluded from our analysis. responses faster than 200ms were assumed to reflect insufficient processing of the stimulus. responses slower than 2,500ms were not taken to exclusively involve the early moments of processing focus relevant to our research question. these exclusion criteria resulted in less than 10% data loss across conditions. bayesian mixed-effects models were used to analyze accuracy and log transformed response time data via the brms package (bürkner 2017) in the r software environment (r core team 2023).5 no divergent chains were observed and all models converged with r̂ ≈ 1 and sufficient 2in english, the l+h* pitch accent is associated with both the presence of focus and sentence final nuclear pitch accent (büring 2016). given that focus always occurred sentence final, our stimuli are technically ambiguous with respect to prosody. however, the presence of the focus particle only and the givenness of the surrounding non-focused material eliminated any possible interpretive ambiguity. see muxica & harris (to appear) for further discussion. 3pairwise differences between each of the probe word conditions were evaluated with a bayesian t-test and no reliable differences in any measure was observed (89% cri, bf<1) for each comparison. 4for the lsa norming, pairwise differences between each of the probe word conditions were evaluated with a bayesian t-test. as intended, associate probe words were found to have a higher cosine similarity to the focused word than non-associate probe words (med=0.5, cri89%=[0.45, 0.54], bf>1,000). the results of the norming study were integrated into our analysis. we fit bayesian mixed-effects models for both response time and accuracy using semantic similarity ratings as a random effect, specifically the log transformed difference between the associate and non-associate probe. in all cases, this addition neither improved model fit nor changed the qualitative pattern of results. we have included these models on the osf repository for this paper (https://osf.io/kmdu3/?view_only= 92f204af070c433d9db6f61ab7c30123). 5comparable frequentist mixed-effects models were also fit using the lme4 package (bates 2010). in all cases, the results were qualitatively the same and thus we do not report these models in the main text. we have made these proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 268 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ effective sample sizes (ess) for each parameter. posterior predictive checks graphically confirmed that the model was an appropriate fit of the response variable. 4.3. internet pilot. a pilot experiment was conducted over the internet using a subset of the materials. this not only allowed us to test the design and validate the central effects in a different setting, but also to generate an informative prior for use in the bayesian analysis of the main in-person experiment. 4.3.1. participants. forty-five self-reported native english speaking undergraduates were recruited from the university of california los angeles psychology department subject pool and given course credit in exchange for participation. 4.3.2. procedure. four lists of 12 items (selected from the set of 28 critical items) were created in a counterbalanced design. eighteen items from the set of 56 filler items were added to these lists yielding 30 trials per list. pcibex was used to host the experiment (zehr & schwarz 2018). on each trial, subjects were presented with a central fixation cross while the audio dialogue played. immediately after the audio completed, subjects were presented with a written probe word in the center of the screen. subjects were instructed to provide their response to the probe word as quickly as possible without sacrificing accuracy. there was no explicit timeout for long responses. after each trial, subjects were given the opportunity to take a self-paced break. the pilot took approximately 12 minutes to complete on average. 4.3.3. results. in the one-alt context, probe task accuracy task was higher on average for the associate (m=94%, se=2) than the non-associate (m=78%, se=4) probe word. the same pattern held for the two-alt context; probe task accuracy was higher on average for the associate (m=95%, se=2) than the non-associate (m=86%, se=3) probe word. however, according to a bayesian logistic regression model, there was no evidence for an effect of probe word (med=0.898, cri89%=[-2.306, 0.483]), context (med=-0.134, cri89%=[-0.77, 0.511]), or an interaction between the two (med=-0.188, cri89%=[-0.844, 0.425]). log transformed response times were subjected to a bayesian linear mixed effects model. main effects are presented first, followed by the interaction. as shown in table 1 in the next section, the model indicated that, across contexts, the non-associate probe word (m=1189ms, se=45) yielded slower response times than the associate probe word (m=1058ms, se=38). there was no evidence that one-alt and two-alt contexts elicited different reaction times. the magnitude of the response time difference was much larger in the one-alt context (131ms) than the two-alt context (54ms), which was in line with an interactive effect between probe words in the two contexts. although there was no evidence of an interaction in the main model, investigation of the estimated marginal means of the model was consistent with the predicted interactive effect. response times to the associate probe were faster than the non-associate probe in the one-alt context (med=-0.020, hpd=[-0.035, -0.004]), but there was no difference between probe words in the two-alt context (med=-0.009, hpd=[-0.023, 0.005]). 4.3.4. discussion. despite its limited the power, the pilot provided preliminary support of a key prediction of the immediate-access model. response times to a probe word depended on whether models available on the osf repository for this paper. proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 269 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the word was presented as a contextually-relevant alternative, rather than on its lexical association with the word in focus. crucially, the effect was observed immediately after presentation of focus. under the two-stage model, response times to a given probe word should solely depend upon semantic priming at early stages of focus interpretation, rather than contextually-determined alternatives status. at this point, it is not entirely clear how such a model could explain the immediate advantage for contextually-determined alternatives. however, evidence for the crucial interaction was only observed in the marginal means, and we now turn to the main, in-person experiment with increased power and a more controlled experimental setting. 4.4. in-person experiment. 4.4.1. participants. fifty-one self-reported native english speaking undergraduates were recruited from the same population population as the pilot and given course credit in exchange for participation. 4.4.2. procedure. the experiment was presented using linger (rhode 2001) and was hosted on a linux desktop computer in a sound-attenuated booth. subjects listened to the audio through seinheiser hd280 pro wired headphones and provided all responses using a ps/2 keyboard. on each trial, subjects were presented with a central fixation cross while the audio dialogue played. immediately after the audio completed, a written probe word was presented in the center of the screen. subjects were instructed to provide this response as quickly as possible without sacrificing accuracy, but there was no explicit timeout for long responses. in addition, multiple choice comprehension questions were presented after a third of the trials. subjects were instructed to prioritize accuracy over speed in answering these questions. as in the pilot, subjects were given the opportunity to take a self-paced break after each trial. the experiment took approximately 30 minutes to complete on average. 4.4.3. results. accuracy results of the main study was comparable to that of the pilot study. in the one-alt context, probe task accuracy task was higher on average for the associate (m=92%, se=1) than the non-associate (m=73%, se=3) probe word. in the two-alt context, probe task accuracy was also higher on average for the associate (m=90%, se=2) than the non-associate (m=94%, se=1) probe word. according to a bayesian logistic regression model, there was no reliable evidence indicating an effect of probe word (med=-0.58, cri89%=[-1.15, 0.01], bf=0.89) or an effect of context (med=-0.34, cri89%=[-0.68, 0.01], bf=1.77). however, there was strong evidence for an interaction between probe word and context (med=-0.66, cri89%=[-1.00, -0.33], bf=51.29). the interaction was further supported by the estimated marginal means of the model, which indicated that the accuracy for the non-associate probe was lower than the associate probe in the one-alt context (med=2.48, hpd=[0.81, 4.24]), but did not differ in the two-alt context (med=-0.163, hpd=[-1.79, 1.39]). log transformed response times were subjected to a bayesian linear mixed effects model. the model, summarized in table 1, indicated that, across contexts, the non-associate probe word yielded slower response times than the associate probe word (non-associate: m=1163ms, se=39; associate: m=1014ms, se=24; bf>100). there was no evidence that response times differed between the one-alt and the two-alt context (bf=0.662). however, there was strong evidence in proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 270 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ favor of the crucial interaction effect between probe word and context predicted by the immediateaccess model (bf>100). the interaction was further supported by the estimated marginal means of the model. response times to the associate probe were faster (149ms) than the non-associate probe in the one-alt context (med=-0.024, hpd=[-0.032, -0.015]), but the difference between probe words in the two-alt context (45ms) was not reliable (med=-0.006, hpd=[-0.014, 0.001]). pilot in-person parameter median 89% cri median 89% cri bf intercept 1.941 [1.930, 1.952] 1.938 [1.931, 1.945] na non-associate vs. associate 0.007 [0.002, 0.012] 0.007 [0.005, 0.010] >100 one-alt vs. two-alt 0.000 [-0.004, 0.004] 0.001 [-0.001, 0.002] 0.662 probe word x context 0.003 [-0.001, 0.006] 0.004 [0.002, 0.006] >100 table 1: results for pilot and in-person studies from bayesian linear mixed effects regression model on log response times with maximal random effect structures and sum-coded predictors. in the pilot study, uninformative (flat) priors were used and so no bayes factor could be computed. the model was run with 5,000 iterations and a 1,000 iteration warm up, and converged with r̂ = 1 and at least an 2,500 ess per parameter. in the in-person study, informative priors from the pilot study were used. the model was run with 12,000 iterations and a 2,000 iteration warm up, and converged with r̂ = 1 and an ess ≥4,000 per parameter. bayes factor (bf) was computed over a null point estimate using the savage-dickey density ratio. a central prediction of the immediate-access model is that context can modulate whether a semantically unrelated word is immediately accessed as an alternative of a word in focus. the interaction we observed in response times is highly compatible with this prediction. the cost observed for non-associate over associate probes in the one-alt context was absent in the twoalt context. however, the strongest manifestation of this prediction is in a “3 against 1” pattern, in which the one-alt non-associate condition elicits longer reaction times than the other three conditions. while the crucial penalty for non-associate over associate probes was evident in the results, the precise pattern is less clear. to explore the interaction in more detail, we fit an exploratory model with trial half as an interactive predictor and random effect, which yielded a better fit than the original. in the trial half model, the overall qualitative pattern of results did not change: non-associate probes yielded slower response times than associate probes (bf>100), response times did not differ between contexts (bf=0.686), and there was strong evidence in favor of an interaction between probe word and context (bf>100). in addition, response times in the first half of trials differed from those in the second half (bf>1,000). despite an impressionistic difference between trial halves, a three-way interaction between context, probe type, and trial order was not supported by the model. figure 1b depicts the response times by condition along trial half. unsurprisingly, response times in the second half of trials (m=987ms, se=17) were much faster than in the first half of trials (m=1178ms, se=18) overall, suggesting that subjects became more adept at the task over time. more importantly, the predicted interaction was observed in both halves. however, the patproceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 271 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ tern of interaction in the first half of trials more closely aligns with the precise three against one pattern predicted by the immediate-access model, in which the one-alt non-associate condition elicited slower responses than the other three conditions. in the second half, there was an additional response time advantage for the one-alt associate condition. we further speculate on the interpretation of the differential effects across conditions in the discussion. all trials one−alt two−alt 900 1000 1100 1200 1300 context r es po ns e t im e (m s) probe word associate non−associate mean response time by condition (a) collapsing across trial order first half second half one−alt two−alt one−alt two−alt 900 1000 1100 1200 1300 context r es po ns e t im e (m s) probe word associate non−associate mean response time by condition (b) splitting along trial order figure 1: main study. mean response time results by condition. error bars indicate standard errors. 4.4.4. discussion. the results of the main experiment provide strong evidence in favor of the immediate-access model, largely replicating the pilot study. the most important finding is that the contextual status of a given probe word as a focus alternative determined response speed across conditions. specifically, the associate the non-associate probes elicited similar response times in the two-alt contexts, when both probe types were presented as contextually-relevant focus alternatives. however, in the one-alt context, where only the associate probe was a relevant focus alternative, we found a response time penalty for the non-associate probe word. crucially, this advantage for alternatives over non-alternatives manifested immediately after the word in focus was encountered. further, our results cannot be explained in terms of semantic priming from the element in focus. in particular, the response time advantage for the non-associate probe word tracked our manipulation of the discourse context rather than semantic association with the element in focus. only the immediate-access model directly predicts that context-sensitivity should manifest immediately. in contrast, the two-stage model predicts that response speed should solely depend upon semantic association at early stages. as mentioned, an interaction was observed in both halves of the study. however, the character of the interaction differed slightly in later trials. though there are many possible explanations for effects of exposure, we speculate that subjects developed a strategy to better predict what the upcoming focus will contrast with from the context and the target sentence. in each condition, the target sentence presented just one new content word (e.g., violin), while the remainder of the sentence frame presented discourse-given information (no, he only brought the . . . ). in the twoalt context, subjects may have been able to predict that the focused word would relate to one of the alternatives (e.g., guitar or pizza), but would not have enough information to prioritize one over the other. in the one-alt context, only the associate (guitar) appeared previously with the sentence proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 272 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ frame, thus allowing the subject to make a more fine-grained prediction that the focused word would contrast with the associate, potentially re-activating it in memory. in other words, subjects learned to make more specific predictions for the associate in the one-alt condition, which gave them an advantage when the prediction was confirmed. in support of this interpretation, subjects became more accurate just in the one-alt associate condition in the second half of the experiment (m=95%, se=2) compared to the first half (m=89%, se=2). in all other conditions, accuracy decreased in the second half. this line of reasoning is supported by the literature. there is considerable evidence that readers and listeners continuously form predictions during comprehension (e.g., staub 2015, kuperberg & jaeger 2016; for review) and that they exhibit processing difficulties when their predictions fail to be validated (e.g. rich & harris 2021, 2023; among many others). it is likely that many sources of information contribute to prediction formation and that focus, and perhaps information structure more generally, is another strong factor. in addition, subjects adapt their processing strategies when presented with mismatching contrastive accent (roettger & franke 2019, nakamura et al. 2019) over the course of an experiment. far from being an experimental nuisance, such strategies can be construed as rational adaptations to the task and may even reflect fundamental aspects of language processing. whatever the case may be, the effect of trial order observed here raises interesting general questions about the potential relationship between focus and predictability, as well as how subjects may strategically adapt to the structure of the experiment. 5. conclusion. across two cross-modal probe recognition experiments, alternative status of a given probe word was found to immediately modulate response speed. the results strongly support the view that an unrelated word can be immediately construed as a focus alternative depending on the context. we take these results to strongly support the immediate construction of focus alternatives by context, as predicted by the immediate-access model. the results provide further evidence against versions of a two stage model in which focus alternatives are initially determined by contextinsensitive lexical association with the word in focus. in all, we believe that the results showing the immediate availability of a contextually determined alternative set strongly coheres with the anaphoric component inherent in calculating the effect of focus on a sentence in the context of utterance (rooth 1992). indeed, constructing contextually-relevant focus alternatives might well constitute a grammatically mandatory operation that cannot be delayed during interpretation (frazier 1999). we leave this, and other questions of focus interpretation and the architecture of the language processing system, to future research. references balota, david a, melvin j yap, keith a hutchison, michael j cortese, brett kessler, bjorn loftis, james h neely, douglas l nelson, greg b simpson & rebecca treiman. 2007. the english lexicon project. behavior research methods 39. 445–459. bates, douglas m. 2010. lme4: mixed-effects modeling with r. beaver, david i & brady z clark. 2009. sense and sensitivity: how focus determines meaning. john wiley & sons. proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 273 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ braun, bettina & lara tagliapietra. 2010. the role of contrastive intonation contours in the retrieval of contextual alternatives. language and cognitive processes 25(7-9). 1024–1043. brysbaert, marc & boris new. 2009. moving beyond kučera and francis: a critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for american english. behavior research methods 41(4). 977–990. büring, daniel. 2016. intonation and meaning. oxford university press. bürkner, paul-christian. 2017. brms: an r package for bayesian multilevel models using stan. journal of statistical software 80(1). 1–28. frazier, lyn. 1999. on sentence interpretation, vol. 22. springer science & business media. gotzner, nicole & katharina spalek. 2019. the life and times of focus alternatives: tracing the activation of alternatives to a focused constituent in language comprehension. language and linguistics compass 13(2). e12310. gotzner, nicole, isabell wartenburger & katharina spalek. 2016. the impact of focus particles on the recognition and rejection of contrastive alternatives. language and cognition 8. 59 – 95. hoeks, morwenna, maziar toosarvandani & amanda rysling. 2023. processing of linguistic focus depends on contrastive alternatives. journal of memory and language 132. 104444. husband, matthew & fernanda ferreira. 2016. the role of selection in the comprehension of focus alternatives. language, cognition and neuroscience 31(2). 217–235. jördens, kim a., nicole gotzner & kathrina spalek. 2020. the role of non-categorical relations in establishing focus alternative sets. language and cognition 12(4). 729–754. krifka, manfred. 1992. a framework for focus-sensitive quantification. semantics and linguistic theory 2. 215–236. kuperberg, gina r & t florian jaeger. 2016. what do we mean by prediction in language comprehension? language, cognition and neuroscience 31(1). 32–59. lacina, radim, patrick sturt & nicole gotzner. 2023. the comprehension of broad focus: probing alternatives to verb phrases. landauer, thomas k & susan t dumais. 1997. a solution to plato’s problem: the latent semantic analysis theory of acquisition, induction, and representation of knowledge. psychological review 104(2). 211. muxica, christian j. & jesse a. harris. to appear. constructing alternatives: evidence for the early availability of contextually relevant focus alternatives. in nicole gotzner, jesse harris, richard breheny & yael sharvit (eds.), alternatives in grammar and cognition, palgrave. nakamura, chie, jesse a. harris & sun-ah jun. 2019. listeners’ beliefs influence prosodic adaptation: anticipatory use of contrastive accent during visual search. in sasha calhoun, paola escudero, marija tabain & paul warren. (eds.), proceeding of the 19th international congress of phonetic sciences, 447–451. pierrehumbert, janet & julia hirschberg. 1990. the meaning of intonational contours in the interpretation of discourse. in intentions in communication, 271–311. mit press. r core team. 2023. r: a language and environment for statistical computing. r foundation for statistical computing vienna, austria. https://www.r-project.org/. rhode, doug. 2001. linger. a flexible platform for language processing experiments . rich, stephanie & jesse a. harris. 2021. unexpected guests: when disconfirmed predictions proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 274 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ linger. in the proceedings of the 43rd annual meeting of the cognitive science society, 2246– 2252. vienna, austria. rich, stephanie & jesse a harris. 2023. global expectations mediate local constraint: evidence from concessive structures. language, cognition and neuroscience 38(3). 302–327. roettger, timo b & michael franke. 2019. evidential strength of intonational cues and rational adaptation to (un-) reliable intonation. cognitive science 43(7). e12745. rooth, mats. 1992. a theory of focus interpretation. natural language semantics 1(1). 75–116. staub, adrian. 2015. the effect of lexical predictability on eye movements in reading: critical review and theoretical interpretation. language and linguistics compass 9(8). 311–327. zehr, jeremy & florian schwarz. 2018. penncontroller for internet based experiments (ibex). proceedings of elm 3: 262-275, 2025 christian muxica and jesse a. harris: constructing focus alternatives from context and the limits of semantic priming. 275 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ analyzing naturally-sourced questions under discussion karl mulligan & kyle rawlins* abstract. the question under discussion (qud) framework of discourse has been a highly influential theoretical device in many accounts of various pragmatic phenomena, yet there has been comparatively little work assessing the extent to which the qud can be reliably inferred from naturalistic contexts. in this paper, we focus primarily on measuring the variability across individuals in qud inference, while also verifying other related, commonly held assumptions about qud theory. to this end, we collect quds from many theoretically naive subjects tasked with processing a radio interview utterance by utterance. we consider various analyses designed to address the problem of measuring question similarity. overall, we find that there exists moderate variability among subjects, consistent with possibly the insufficiency of context in determining qud, or possibly also the simultaneous coexistence of multiple valid quds. to more adequately tease apart these possibilities, we also propose additional analyses for addressing the issue of question identity. keywords. question under discussion; pragmatics; discourse structure; question similarity 1. introduction. one of the most influential paradigms in pragmatics is the question under discussion framework, which analyzes discourse as a process guided by addressing implicit questions (van kuppevelt 1995, roberts 1996/2012). the distinguishing feature of this framework is that contextual relevance is modeled using natural language questions. one advantage of this is that the framework is able to profit from an existing rich and successful tradition in formal semantics of analyzing questions as sets of alternatives (hamblin 1973). but another advantage of using questions is the fact that natural language questions are, of course, easily understood and generated by non-specialists. we exploit this latter advantage to investigate our central question: given the same access to discourse context, do normal language users reliably infer the same implicit qud? underlying most formal accounts of qud-sensitive phenomena is the assumption that the immediate qud is accessible, or at least inferrable, for all discourse participants at each stage of conversation. while this may be a reasonable assumption, there is limited work so far exploring the extent to which this underlying assumption holds in natural discourse. we also have only limited evidence that the kinds of quds posited to account for various pragmatic inferences are indeed observed “in the wild” within typical conversation. we are thus motivated to collect a variety of quds from naturalistic discourse, using natural language questions produced by theoretically naive subjects. given a sense of what such naturally-sourced quds are like, this may help determine the ecological validity and generality of accounts which depend on quds. in this work, we break down the fundamental issue of qud accessibility into three more specific assumptions about quds commonly found in the literature: *we would like to thank our anonymous reviewers, as well as audiences at macsim, elm, and scil for their feedback. authors: karl mulligan, johns hopkins university (karl.mulligan@jhu.edu) & kyle rawlins, johns hopkins university (kgr@jhu.edu). proceedings of elm 3: 250-261, 2025 c©2025 karl mulligan and kyle rawlins published by the lsa with permission of the author(s) under a cc by license. 250 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ • singularity: there is a single immediate qud at any given turn in discourse. • q-to-qud: an explicitly asked question becomes the new immediate qud. • q-a congruence: the answer to the immediate qud is contained within the asserted utterance, with the wh-word congruent, or corresponding, to the focused constituent (rooth 1992). to test these assumptions, we assess to what extent ordinary language users are able to produce and agree upon the immediate qud while processing naturalistic dialogue. building on data collected by mulligan & rawlins (2024), we find that there is persistent variability across participants in their inferred quds, though the data are also consistent with frequent, non-random pockets of agreement among participants, a finding which calls singularity into question but stops short of disproving it. we also find results suggesting that q-to-qud may be too strong of an assumption. for the last assumption of q-a congruence, however, we find that our subject-generated quds are largely congruent to selected answer spans, which are usually grammatical constituents. overall, we find that some contextually-constrained notion of qud is accessible and shared across language users, but we are limited by our methods and analyses in our ability to narrow down the source of this variability. 2. background. although there are various theories of discourse that depend heavily on some notion of implicit question (van kuppevelt 1995, ginzburg 1996), this work primarily focuses on the model described in roberts (1996/2012). for roberts, discourse is likened to a game, in which the various moves (utterances) are recorded on a public scoreboard. in addition to the common ground, and the sets of explicitly uttered assertions and questions, there is another, generally implicit component, the question under discussion stack. a move in discourse is deemed relevant if it addresses the topmost item on this stack, the immediate qud. in principle, the next qud can be inferred as a function of interlocutor goals and the information in this discourse data structure: prior conversational moves, shared common ground assumptions, and existing quds on the stack (cooper et al. 2000, velleman & beaver 2016). however, the exact process by which this inference is assumed to take place is far from fully understood. qud annotations are thus a potentially valuable resource for this effort. there exist several works, from both linguistics and natural language processing, that attempt to source implicit questions from natural language. the approaches vary along several dimensions. some works, like de kuthy et al. (2018), involve detailed, hierarchical annotation paradigms that hew closely to the formalism described by roberts (1996/2012), thus requiring trained annotators familiar with qud theory. other works use crowdsourced participants (westera et al. 2020, pyatkin et al. 2020, wu et al. 2023), though often using more general notions of implicit question agnostic to specific structural assumptions. common to all of the approaches mentioned above is the fact that the source material is monologue, and typically written rather than spoken. moreover, prior efforts in this area are generally limited to one or two annotators per item, making quantifying variability difficult, though de kuthy et al. (2018) among others report frequent qualitative discrepancies in annotation decisions. 3. methods. in this work, we sourced our quds from dialogue, and in order to quantify variability in qud inference, we crowdsourced a large number of implicit questions for each utterance. proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 251 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3.1. materials. for our conversation data, we picked 10 two-party interviews from the interview dataset (majumder et al. 2020), a collection of national public radio interviews in american english. while perhaps not as spontaneous as some other genres of speech, these interviews contain few disfluencies and are of consistently high quality, facilitating sentence-by-sentence annotation. interviews were also generally conversational in tone, making them suitable for studying the processing of naturalistic discourse. for consistency, we chose interviews with between 29 and 32 sentences, at least 5 of which were explicitly asked questions. 3.2. participants. for each interview, we recruited 10 native english speakers per episode via profilic, for a total of 100 unique sets of qud annotations. this high number of participants per episode allows us to get a highly calibrated measure of variability in qud selection across participants at the utterance level. participants took approximately 20 minutes to complete each interview, and were compensated $15 per hour for their work. 3.3. procedure. participants were presented the interview in a moving two-sentence window in order to simulate real-time, incremental revelation of context, following westera et al. (2020). the first of the two sentences was the context sentence, while the second was the target sentence, always indicated in bold. for each sentence pair, they were instructed to do two things: first, to write a question that could be answered by the target sentence; and second, to select the shortest contiguous span of words from the target sentence which best answers the question they just wrote. for utterances which did not serve as answers to a clear question (such as commands, requests, greetings, and other non-declarative sentences), participants were instructed to check a box labeled “no clear question” and, instead of giving a question–answer pair, to explain why there was no clear question being addressed. as a complimentary task testing the q-to-qud assumption, we also masked explicitly asked questions to see whether subjects would produce quds for the subsequent utterance resembling the asked question. for each literal question asked in the interview, we replaced it with “[question masked]” while keeping the utterance before it intact and visible as the context sentence. these trials were not any different otherwise from regular trials. 3.4. evaluation. in order to measure the variability in inferred quds across subjects, we employed three question similarity metrics, each with its own advantages and disadvantages. the first of these is word edit distance, which is a sum of the minimum number of edits (insertions, deletions, or substitutions of words) needed to transform one sentence into another. under the assumption that similar questions can be addressed by similar answers, we we also use edit distance to measure similarity between answer spans. this makes answer edit distance a useful proxy for question similarity (indeed, we find that these measures are correlated), and also more reliable since all answers for a given trial are subsets of the same utterance. the second metric is bertscore, a general-purpose neural sentence similarity metric (zhang et al. 2020). bertscore uses a large transformer language model to encode each question into a high-dimensional vector space, and then measures the cosine similarity of these encodings; higher values indicate greater semantic similarity. unlike edit distance, neural metrics like bertscore are less sensitive to variation in the surface input, and are thus more suitable for capturing paraphrases. here we used a rescaled version of deberta, the default configuration suggested by the proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 252 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: response interface from mulligan & rawlins (2024). the question box (q:) is a free response prompt. the answer box (a:) can only be filled by selecting from the target sentence, ensuring that a subject’s written qud is addressable using a contiguous span of the target utterance. bertscore authors. the third metric is wh-word agreement which we define as 1 if both questions share the same wh-word (or in the case of polar questions, the same first word auxiliary), and 0 otherwise. while simple, this metric is useful for coarsely differentiating questions seeking distinct types of information. to illustrate the use of these metrics on real data, we apply each metric to pairs of the following collected quds for the utterance in figure 1, repeated from mulligan & rawlins (2024). the values for each metric are displayed in table 1. (1) who else had been watching the radar? [one of my graduate students] (2) who saw the occurrence and effects on the radar? [my graduate student] (3) where are the clouds coming from? [southwest about five miles] mulligan & rawlins (2024) find that all of these metrics are moderately correlated with one another. the results going forward will mostly use bertscore to measure qud variability because of its granularity and robustness to synonymy. 4. results. 4.1. qud variability. we turn now to an assessment of the first of our three assumptions, the singularity of qud at a given stage of discourse. if there is at most a single qud at any given time, we would expect less pairwise variability among subjects, since the inference of the current proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 253 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ metric ((1),(2)) ((2),(3)) ((1),(3)) word edit distance (a) 3 4 5 word edit distance (q) 6 8 7 bertscore (q) 0.41 0.14 0.12 wh-word agreement (q) 1 0 0 table 1: a comparison of similarity metrics for both answer spans (a) and questions (q) on collected quds. in this example, we wish to capture the qualitative intuition that (1) and (2) are most similar to each other, and are thus candidates for being qualitatively characterized as “the same” qud. figure 2: each bar represents the average pairwise variability among all ten subjects for a single utterance (n = 241 observations), with items ordered by the mean bertscore. qud should be transparent from discourse context. we find that there is moderate variability among subjects in qud inference, calling a strict notion of singularity into question. to visualize this overall variability, we plot the distribution of the mean bertscore per item, ordered by ascending value, in figure 2. a concave distribution for this plot would indicate high overall support for singularity, since the quds for the majority of utterances would have higher mean pairwise similarity, and the distribution would have more overall weight. however, what we in fact see is a somewhat convex distribution: the majority of trials have a mean pairwise bertscore around 0.3, indicating only moderate similarity across quds. for a qualitative view of this variability, as well as more examples of interpreting bertscore values, we present in the appendix a sample of quds from a low-agreement trial (table 2, mean bertscore of 0.23) as well as a sample from a high-agreement trial (table 3, mean bertscore of 0.49). what is the source of this variability? this is a difficult question to answer given our data, but proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 254 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ there are at least a few possible explanations that can be ruled out. for instance, qud variability does not seem to be correlated with time (i.e., position in the interview). one might suppose that overall qud variability decreases over the course of a dialogue, as discourse context accumulates and constrains the space of relevant questions; alternatively, one might instead expect overall qud variability to increase, as more content and discourse referents are introduced. in fact, neither trend is the case: qud variability does not trend significantly in one direction or the other over time. 4.2. masked question analysis. in order to assess our second commonly encountered assumption, q-to-qud, we compare masked explicitly asked questions to subject-written quds for the utterances following them. q-to-qud says that, under typical circumstances (i.e., for accepted, canonical questions) an explicitly asked question becomes the new qud. we find that across the board, for post-masked question trials, subject quds are consistently less similar to the masked question (mean bertscore of 0.172) than they are to one another (mean bertscore of 0.369). this finding seems to suggest that q-to-qud should not be taken for granted as a default assumption, albeit with some caveats. there may be unaccounted for properties of explicitly asked questions which make them harder to reconstruct than ordinary quds; for instance, interlocutors may be more likely to ask an explicit question during a topic shift or when seeking information about an as-yet unmentioned entity. it may also be the case that our setup encourages subjects to write questions which reuse words from the target utterance, thereby artificially positively biasing inter-subject comparisons over masked question comparisons. 4.3. constituency analysis. for the assessment of our third and final assumption, q-a congruence, we perform a constituency analysis of the selected answer spans to see whether they are syntactically congruent to subject-written quds. a possible concern over sourcing quds from non-specialists, as we have here, is that the resulting questions and their answer spans may not reflect essential formal properties of quds, such as the connection to focus and designated alternative sets. to determine whether our sourced quds obey the property of q-a congruence, we perform a constituency analysis on all utterances and answer spans to see whether subjects generally pick grammatical constituents as answers. upon parsing these structures, we can then check whether the syntactic categories of constituents correspond to the appropriate wh-word. to parse the utterances, we used the constituent-treelib library, a neural constituent parser based on the berkeley neural parser and spacy (halvani 2024). for each utterance in each interview, we perform a constituent parse and recursively obtain a set of all constituents in the sentence labeled by syntactic category. we find that just under half of the answer spans are constituents, when directly compared with the outputs of our constituency parser. this is less than expected if we would like to claim that q-a congruence holds in our data. however, due to limitations of using an automatic parser, we notice that there are many false negatives. for instance, an answer span might not be counted as a constituent if the span fails to include an optional adjunct clause found in the original sentence, even if it is perfectly capable of standing alone as a syntactically complete noun phrase. to remedy this problem, we define the notion of constituent leaf coverage for a given answer span, sentence, and constituent parse: proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 255 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ constituent leaf coverage = # of answer tokens # of leaves in smallest common parent subtree this notion gives us a more flexible, gradient measure of constituency: spans which are “nearconstituents” are defined as those which are syntactically close to a phrasal constituent (i.e., they are contained by a common parent subtree of nearly the same size). for these answer spans, the ratio will be close to 1; for answer spans which perfectly match a constituent, this measure will be exactly 1. on the other hand, for answer spans which cross large syntactic boundaries, this ratio will be lower. in figure 3, we see that while half of answer spans are constituents, the distribution of leaf coverage values for non-constituents skews toward 1, indicating that even non-constituent spans are likely to be syntactically informative. figure 3: distribution of constituent leaf coverage for all answer spans. to see whether q-a congruence holds in our data, we examine which constituent phrase types pattern with which kinds of wh-words (or auxiliaries, for polar questions). for these analyses, we limit our answer spans to only perfect constituents. about 70 percent of our quds are wh-questions, while 21 percent are polar questions (we set aside the remaining 9 percent of questions which contain neither a wh-word nor an auxiliary as their first word). in figure 4a, we show a breakdown of answer span syntactic phrasal categories, separated by qud question type. polar questions are shown to be most often answered by entire sentences, while wh-questions are answered by a greater variety of phrase types depending on the particular wh-word. figure 4b reveals how the distribution of phrase type varies by wh-word. we see that quds containing wh-words which mostly pick out sets of entities, such as who, what, and which, are largely composed of noun phrases; words like where and when contain a noticeably higher relative share of prepositional phrases; and words like how and why are mostly answered using full sentences. taken together, the results of this analysis indicate support q-a congruence in our sourced quds. 5. discussion. our results show that, across participants, there exists moderate variability in inferring the current qud in naturalistic discourse. based on our large sample of quds per utterance, we are able to find qualitative pockets of agreement, with our question similarity metrics reflecting levels of agreement above random, though still far from universal agreement. from these analyses alone we are limited in our ability to determine the source of the variability. still, it is worth discussing the possible causes of this variability and how these causes might be uncovered in future work. proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 256 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (a) (b) figure 4: (a) distribution of constituent types for wh-questions and polar questions. (b) proportion of answer span constituent types by wh-word. the data is consistent with at least two general possibilities: there may be uncertainty about the current qud (assuming there is only one intended or “true” qud) and discourse context is simply insufficient in narrowing it down; or there may be an inherent multiplicity of discourserelevant quds, in which case singularity is too strict of an assumption and must be relaxed. of course, these possibilities are not mutually exclusive. there may be multiple active quds, but with varying degrees of discourse relevance and accessibility; the perceived levels of activation for these quds may vary based on individual differences and discourse goals. indeed, this sort of multiplicity is consistent with a view of qud tracking in which questions pushed onto the stack earlier in discourse are still able to be directly addressed, something which can be captured using a hierarchical model. to address the issue of whether there are multiple simultaneous quds, clustering may be a useful, unsupervised method of analysis. agglomerative (or bottom-up) hierarchical clustering starts with each data point as its own cluster, and at each step, the two closest (most similar) clusters join, resulting in a a hierarchically-ordered grouping of similar data points. while we do not perform any extensive clustering analysis in this paper, we would like to share some preliminary clustering results as a way of illustrating the potential of this method to more definitively test the singularity assumption. in figure 5, we show two dendrograms obtained by performing agglomerative clustering on each set of quds from one of our interviews, using ward linkage and bertscore as a euclidean distance metric. in the first dendrogram, all subjects write very similar quds, similar enough that they are counted as a single cluster. in the second, there is more variability among quds: here the algorithm instead groups questions into several distinct, but internally similar clusters. for this latter set of quds, this analysis could be proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 257 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (a) single cluster (b) multiple clusters figure 5: (a) highly similar quds arranged into a single cluster. (b) dissimilar quds grouped into separate, internally similar clusters. used to argue that, for that turn in discourse, there are potentially multiple contextually salient questions being addressed. however, there are limitations to this method, too. while the ability to group questions into clusters is appealingly intuitive, clustering alone is not able to tell us whether the variability in inference arises because of inherent multiplicity or uncertainty — though further analysis of each cluster can give us better clues than depending on aggregate measures of variability. also, as with many unsupervised methods, the number of clusters produced is highly dependent on hyperparameters like the linkage criterion and distance threshold, which prevents us from making any definitive claims about the exact number of quds being addressed. we leave overcoming these and other analytical challenges to future work. 6. conclusion. overall, we find that ordinary language users generally make similar inferences about the current question under discussion, but there is considerable variability among subjects in the quds they produce. we interpret this data in light of three commonly held assumptions proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 258 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ about quds in the literature: the singularity of the current qud, q-to-qud, and q-a congruence. given the persistent variability, we believe that a realistic treatment of naturalistic discourse requires the relaxing our first assumption, singularity, though we stop short rejecting it outright. while some of this variability may be due to inherent multiplicity of contextually felicitous quds, some may be due to subject uncertainty or subject differences in how they integrate discourse context, and some may be due to limitations of our question similarity metrics. our data suggest that our second assumption, q-to-qud should also be relaxed, though more work on the nature and incidence of explicitly asked questions is needed to fully validate this suggestion. lastly, we find that, despite a lack of explicit linguistic training, our subjects write quds which largely follow the assumption of q-a congruence. through further collection and analysis of naturally-sourced quds, we hope to gain a better understanding of how qud inference plays out in naturalistic discourse. references cooper, robin, staffan larsson, elisabeth en-gdahl & stina ericsson. 2000. accommodating questions and the nature of qud . de kuthy, kordula, nils reiter & arndt riester. 2018. qud-based annotation of discourse structure and information structure: tool and evaluation. in proceedings of the eleventh international conference on language resources and evaluation (lrec 2018), miyazaki, japan: european language resources association (elra). ginzburg, jonathan. 1996. dynamics and the semantics of dialogue. in jerry seligman, dag westerståhl & lawrence cavedon (eds.), logic, language, and computation, vol. 1 csli lecture notes, 221–237. stanford, california: csli publications. halvani, oren. 2024. constituent treelib a lightweight python library for constructing, processing, and visualizing constituent trees. 10.5281/zenodo.10951644. https:// github.com/halvani/constituent-treelib. hamblin, c. l. 1973. questions in montague english. foundations of language 10(1). 41–53. majumder, bodhisattwa prasad, shuyang li, jianmo ni & julian mcauley. 2020. interview: large-scale modeling of media dialog with discourse patterns and knowledge grounding. in proceedings of the 2020 conference on empirical methods in natural language processing (emnlp), 8129–8141. online: association for computational linguistics. 10.18653/v1/2020.emnlp-main.653. mulligan, karl & kyle rawlins. 2024. identifying questions under discussion in naturalistic discourse. society for computation in linguistics 7(1). 357–361. 10.7275/scil.2228. pyatkin, valentina, ayal klein, reut tsarfaty & ido dagan. 2020. qadiscourse discourse relations as qa pairs: representation, crowdsourcing and baselines. in proceedings of the 2020 conference on empirical methods in natural language processing (emnlp), 2804–2819. online: association for computational linguistics. 10.18653/v1/2020.emnlp-main.224. roberts, craige. 1996/2012. information structure in discourse: towards an integrated formal theory of pragmatics. semantics and pragmatics 5. 10.3765/sp.5.6. rooth, mats. 1992. a theory of focus interpretation. natural language semantics 1(1). 75–116. van kuppevelt, jan. 1995. discourse structure, topicality and questioning. journal of linguistics proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 259 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 31(1). 109–147. velleman, leah & david beaver. 2016. question-based models of information structure. in caroline féry & shinichiro ishihara (eds.), the oxford handbook of information structure, 0. oxford university press. 10.1093/oxfordhb/9780199642670.013.29. westera, matthijs, laia mayol & hannah rohde. 2020. ted-q: ted talks and the questions they evoke. in proceedings of the 12th language resources and evaluation conference, 1118–1127. marseille, france: european language resources association. wu, yating, william sheffield, kyle mahowald & junyi jessy li. 2023. elaborative simplification as implicit questions under discussion. 10.48550/arxiv.2305.10387. zhang, tianyi, varsha kishore, felix wu, kilian q. weinberger & yoav artzi. 2020. bertscore: evaluating text generation with bert. appendix low question similarity example qud answer span (l.1) what can this hide? [it can be used to deflect any missiles that might be based on radar return and it can hide aircraft.] (l.2) what are one of the uses of chaff in military operations? [it can hide aircraft] (l.3) what is that in the sky? [missiles] (l.4) are there any other military applications? [it can hide aircraft] (l.5) how can chaff be used in other applications? [it can be used to deflect any missiles that might be based on radar return and it can hide aircraft.] (l.6) how would a missile be affected by chaff? [can be used to deflect] utterance: it can be used to deflect any missiles that might be based on radar return and it can hide aircraft. table 2: a sample of low-similarity quds (mean bertscore: 0.23). proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 260 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ high question similarity example qud answer span (h.1) why were pregnant women scared about the zika virus? [their babies might be born with severe birth defects.] (h.2) why are pregnant woman worried about the zika virus? [fear that their babies might be born with severe birth defects.] (h.3) what makes these women think their babies might be born with birth defects? [women who’ve had the zika virus] (h.4) what do the women from brazil and colombia fear? [from brazil and colombia about pregnant women who’ve had the zika virus and fear] (h.5) why are women in fear of the zika virus? [their babies might be born with severe birth defects.] (h.6) what’s concerning to pregnant women in brazil and columbia? [pregnant women who’ve had the zika virus and fear that their babies might be born with severe birth defects.] utterance: this week, we’ve heard stories from brazil and colombia about pregnant women who’ve had the zika virus and fear that their babies might be born with severe birth defects. table 3: a sample of high-similarity quds (mean bertscore: 0.49). proceedings of elm 3: 250-261, 2025 karl mulligan and kyle rawlins: analyzing naturally-sourced questions under discussion. 261 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ indirect discourse as mixed quotation? an experimental investigation sebastian walter* abstract. the results of an experimental rating study are reported suggesting that self-pointing gestures aligned with a third-person pronoun are acceptable in german indirect discourse (id) utterances. following a proposal by ebert & hinterwimmer (2022) for self-pointing gestures in free indirect discourse (fid), self-pointing gestures in id are interpreted as character viewpoint gestures (cvgs) quoted from the matrix subject. crucially, it is argued that in id, a perspective shift to the matrix subject can take place. it is proposed that id is an instance of mixed quotation involving a demonstration (cf. clark & gerrig 1990, davidson 2015) where self-pointing is quoted from the matrix subject’s original utterance. keywords. perspective taking; gesture semantics; indirect discourse; mixed quotation 1. introduction. speakers have several options to report what someone else thought or said. for example, they can do so by giving a verbatim repetition of the reported speaker’s thought or speech, i.e., they directly quote the words the other person uttered. this is commonly referred to as direct discourse (dd). alternatively, they can simply paraphrase the propositional content of the reported speaker’s original thought or utterance. this is called indirect discourse (id) and is standardly assumed to be non-quotational. a third option is what is called free indirect discourse (fid), a way to report someone’s thoughts or speech without any overt marking, which has been argued to be an instance of highly specialized mixed quotation (maier 2015). particularly interesting is the behavior of so-called indexicals in speech reports, an example being the first-person pronoun i. broadly speaking, the reference of an indexical can vary from context to context, which sets them apart from ordinary expressions. coming back to the first-person pronoun i, this means that it always refers to the current speaker. in dd utterances, however, it receives a so-called shifted interpretation, that is, it is not interpreted from the current speaker’s (= reporting speaker’s) point of view, but rather from the reported speaker’s point of view, as first-person pronouns in dd utterances always refer to the reported speaker. in general, all indexicals shift in dd utterances. as will be shown below, the picture is more complex for id and fid because there, not all indexicals have to shift. clark & gerrig (1990) propose that quotations are demonstrational, meaning that they are depictive rather than descriptive. furthermore, in this approach not only words can be quoted, but, for example, also gestures (including facial expressions), intonation, and even non-linguistic behavior. these ideas have been formalized by davidson (2015). in a study investigating maier’s (2015) claim that fid is an instance of mixed quotation with pronouns and tenses being systematically unquoted, ebert & hinterwimmer (2022) found that a self-pointing gesture (= a pointing gesture *acknowledgments: this research was conducted in the dfg-funded project visual and non-visual means of perspective taking in language, which is part of the priority program 2392 (vicom). i thankfully acknowledge the financial support of the german research foundation (dfg). furthermore, i would like to thank cornelia ebert and stefan hinterwimmer for their valuable comments on this work. finally, i thank lennart fritzsche for his help with conducting the study reported in this paper. authors: sebastian walter, goethe university frankfurt (s.walter@em.unifrankfurt.de). proceedings of elm 3: 423-434, 2025 c©2025 sebastian walter published by the lsa with permission of the author(s) under a cc by license. 423 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ toward the speaker’s body) co-referent with the reported speaker is acceptable in fid. they found the same for first-person pronouns in dd utterances. by unifying davidson’s (2015) account of quotation as demonstration with ebert et al.’s (2020) approach to gesture semantics, they analyze the results as follows: self-pointing gestures are interpreted as gestures quoted from the reported speaker. based on the assumption that id is non-quotational, they hypothesized that self-pointing in id is unacceptable. however, they found that self-pointing was also surprisingly acceptable in id utterances. according to them, this suggests that a similar perspective-shifting strategy as in fid is available in id as well. however, it seems to be more constrained and is not necessarily present. in other words, id allows for the option of mixed quotation to be present. the study reported in this paper further investigates this claim. previous research has shown that temporal and local adverbials, which are also indexicals, can shift in id only if the reported speaker’s perspective is prominent in the surrounding discourse (plank 1986, anderson 2019). based on these findings, it was hypothesized that self-pointing is acceptable in id only if the perspective of the reported speaker (= the matrix subject) is prominent, or, in other words, that mixed quotation is available in id if this perspective is prominent. to test this, a rating study was conducted. the results go beyond the initial hypothesis, since they show that self-pointing in id is acceptable regardless of whether the reported speaker’s perspective is prominent on the speech level or not, thus suggesting that id in general allows for mixed quotation. this paper is structured as follows: section 2 provides the relevant theoretical background on perspective in speech (section 2.1), speech-accompanying gestures (section 2.2), and quotation (section 2.3). the experimental study is reported in section 3. section 4 concludes the paper. 2. background. 2.1. perspective in speech. many expressions depend on a perspective in their interpretation. among these expressions are predicates of personal taste ((be) tasty), epithets (that idiot), relational expressions (this, that, left, right), and evaluative expressions (fun). the default for them is to be interpreted from the perspective of the current speaker (harris 2012). however, this default changes in instances of speech reports. as will be shown below, depending on the type of speech report, they can receive what is commonly referred to as a shifted interpretation, i.e., an interpretation not from the current speaker’s perspective, but from the perspective of the speaker whose thoughts or speech is currently being reported. examples of different types of speech reports are given in (1). (1) on her way home, mary heard a song by kendrick lamar that she liked on the radio. a. she thought: “i will buy his new album tomorrow.” b. she thought that she would buy his new album on the following day. c. she would buy his new album tomorrow. (hinterwimmer 2017, p. 284) the utterance given in (1a) is an instance of dd. in dd, the words or thoughts of the reported speaker are directly quoted. therefore, a complete perspective shift toward the reported speaker takes place and all perspective-dependent expressions are evaluated from their perspective. thus, the first-person personal pronoun i refers to mary in (1a) rather than to the current speaker. moreover, temporal adverbials, such as tomorrow in (1a), are also evaluated in the context where the proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 424 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ reported speech event took place, meaning that tomorrow interpreted as referring to the day following mary’s thought. in (1b), where the same thought as in (1a) is rendered as an utterance in id, perspectivedependent expressions behave differently than in dd. id is generally analyzed as a non-quotational report of some other individual’s speech or thoughts. consequently, the systematic shifting behavior of perspective-dependent expressions attested for dd does not fully generalize to id. for example, personal pronouns (at least in german, english, and many other languages) never shift in id (but see, e.g., anand & nevins 2004 for data from other languages where personal pronouns shift in id). this is illustrated by (1b), where a third-person pronoun is needed to refer to mary. using a first-person pronoun would result in a different interpretation of the sentence. by contrast, predicates of personal taste and evaluative expressions receive a shifted interpretation in id. temporal and local deictic expressions are normally interpreted from the reporting speaker’s perspective, thus explaining why on the following day is used to refer to the day following the thinking event in (1b). however, these expressions can also receive a shifted interpretation if the reported speaker’s perspective (i.e., mary’s in (1b)) is prominent in the discourse (plank 1986, anderson 2019). finally, the utterance in (1c) is an instance of fid—a way to report another individual’s thoughts or speech without overt marking. in fid, all perspective-dependent expressions shift toward the reported speaker, the only exceptions being pronouns and tense markings. fid thus shares properties with both id and dd and is often seen as a blend of the two. in formal semantics, the most popular treatment of fid is the so-called double context analysis (e.g., schlenker 2004, eckardt 2014). there is, however, also a proposal by maier (2015), which treats fid as an instance of highly specialized mixed quotation, where a fid utterance is seen as a verbatim quote of an individual’s thoughts or speech with pronouns and tense markings being systematically unquoted. a similar proposal will be made for id in this paper. 2.2. speech-accompanying gestures. gestures are defined as communicative movements of the hands, arms, and even the face which transport a speaker’s thoughts, intentions, or emotions. therefore, gestures that occur alongside speech add meaning to an utterance, or, more precisely, speech and gesture work together to convey a multimodal message (e.g., mcneill 1992, kendon 2004, de ruiter 2000, ebert 2024). the following examples illustrate this: (2) a. i brought a bottle of water to the talk. + big1 b. i brought this bottle to the talk. + pointing to bottle the iconic big-gesture in (2a) contributes size information of the bottle to the verbal part of the utterance, i.e., that it is big. the pointing gesture in (2b), by contrast, selects its referent directly and adds to the verbal part of the utterance that the bottle pointed at is the bottle the speaker brought to the talk. crucially, in the semantic theory modeling the meaning contributions of speech-accompanying gestures by ebert et al. (2020), it is argued that this is the only difference between pointing and co-speech gestures: while iconic gestures refer to an abstract individual, pointing gestures make reference to an individual in the world. besides this, they behave alike 1underlining in examples indicates gesture-speech alignment, small caps are used to indicate a speechaccompanying gesture. proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 425 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ from a semantic point of view in ebert et al.’s (2020) account and are therefore both analyzed as making meaning contributions similar to that of appositives (cf. potts 2005). they thus, unless aligned with a demonstrative as in (2b), make not-at-issue contributions by default (for experimental validation, see ebert et al. 2020), meaning that they cannot be directly denied in discourse and project through semantic operators such as negation. especially this first property has often been ascribed to not-at-issue content not being on the table for discussion (farkas & bruce 2010). the short discourse in (3) illustrates this. (3) a: there is a window in peter’s living room. + round b1: # that’s not true! the window is not round. b2: yes, true, but the window isn’t round. when trying to directly object the gesture’s contribution of a’s utterance as b1 does, this results in infelicity. what one can do instead to target the not-at-issue content conveyed by the gesture is, for example, to assent with the main clause proposition followed by an adversative continuation (cf. diagnostic #1c in tonhauser 2012). this explains why the response of b2 is felicitous. it has been shown in section 2.1 already that perspective plays a decisive role when it comes to the interpretation of many lexical expressions. however, perspective can also be encoded in gesture (mcneill 1992), the most important distinction being the one between so-called character viewpoint gestures (cvgs) and observer viewpoint gestures (ovgs). when performing a cvg, the whole body is often involved in the gesture production as it depicts an event from an internal, firstperson perspective. this contrasts with ovgs where normally only hands and arms are involved in the production of the gesture as it depicts an event from an outside, third-person perspective. the examples given in figure 2a and figure 2b illustrate this. the pictures are taken from a study reported in parrill (2010) where participants described short cartoon clips to a friend they brought to the study. the two persons shown in figure 2 describe the scene shown in figure 1 where a skunk is hopping across a room. the person in figure 2a produces a clear instance of a cvg as they mimic the skunk’s body posture and also its facial expressions, thus adopting a first-person perspective. the person in figure 2b, by contrast, describes the same hopping event by means of an ovg because they only use their right index finger to trace the hopping movement of the skunk. thus, they adopt the external, third-person perspective that is representative of ovgs. from these examples, another interesting can be made: the two types of viewpoint gestures differ with respect to what they typically express. for example, ovgs encode trajectory more often than cvgs do (parrill 2010). another interesting observation is that cvgs have been shown to be more informative than ovgs (beattie & shovelton 2002). several other differences between cvgs and ovgs have been attested. however, they will not be discussed here, as they would exceed the scope of the present paper. 2.3. quotation as demonstration. quotations, such as the dd utterance in (1a), have traditionally been argued to be verbatim repetitions of what someone else said (e.g., quine 1940, davidson 1979). the quoted words were assumed to be mentioned, but not used, suggesting they cannot interact with surrounding linguistic material. more recent research, however, has shifted away from a strict verbatim condition (e.g., davidson 2015), and it is more common to assume that quoted elements are often simultaneously used and mentioned. this is evidenced by examples proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 426 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: cartoon scene of a skunk hopping across a room. taken from parrill (2010) (a) cvg used to depict the skunk in figure 1 (figure 3 in parrill 2010, p. 652) (b) ovg used to depict the skunk in figure 1 (figure 2 in parrill 2010, p. 651) figure 2: examples of a cvg and an ovg to depict the event shown in figure 1 such as the following: (4) stiviano’s lawyer has not denied the part about the gifts, although he says there is not a “peppercorn of a fact” that any fraud was involved. (nyt, may 1, 2014, cited from davidson 2015, p. 483) here, the quoted constituent a peppercorn of a fact is used and mentioned at the same time. it is mentioned because the author of the sentence in (4) indicates by use of the quotation marks that the words quoted are a verbatim repetition of what stiviano’s lawyer said. however, the phrase is also used as it needs to be integrated with the surrounding linguistic material to derive a full proposition. the content of the quotation thus also plays a role in this example. this means that the underlying report paraphrasing the original proposition expressed in the reported speech act can be retrieved when ignoring the quotation marks. the quotation marks add another level of meaning—namely that whatever stands in the quotation marks was part of the original utterance (potts 2005, maier 2015). this is commonly known as mixed quotation. for fid utterances (cf. (1c)), an analysis has been proposed treating it as a highly specialized form of mixed quotation where the original thought or utterance of the reported speaker with pronouns and tense markings being systematically unquoted, thus accounting for their shifting behavior (maier 2015). the unquotation mechanism is conceptualized as a pragmatic process. the quotational structure of (1c) is thus as in (5). square brackets indicate unquotation. proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 427 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (5) on her way home, mary heard a song by kendrick lamar that she liked on the radio. “[she] [would] buy [his] new album tomorrow.” note that the quotation marks and the square brackets are just used to illustrate (un)quotation. normally, they are omitted in fid utterances. before turning to experimental evidence in favor of the mixed-quotational account of fid proposed by maier (2015), let’s briefly introduce the idea that quotation involves demonstration (clark & gerrig 1990). this means that quotations in general serve a depictive rather than a descriptive purpose and are thus iconic in their nature. additionally, the idea that quotation involves demonstration implies that speakers can quote aspects of an utterance relating to its form and not only to its content. (6) and so she said “[whispering] what are we going to do?” (clark & gerrig 1990, p. 487) here, not only words are quoted, but also the manner in which they were uttered, that is, with a whispering voice. the predictions of the demonstrational account to quotation go even further, as evidenced by (7). the be like-construction is also an instance of quotation (e.g., davidson 2015) where special emphasis is put on quoting the form of an utterance. the quotation marks are again only inserted to highlight the quoted part of the utterance and are normally omitted. (7) my cat was like “feed me!” (davidson 2015, p. 485) by virtue of uttering (7), an event is demonstrated, i.e., quoted, where the speaker’s cat behaves in a certain manner leading them to infer that it is hungry. this illustrates that when assuming that quotation is inherently demonstrational, even non-linguistic behavior can be quoted. davidson (2015) proposes a formal analysis for these observations and explicitly claims that also speech-accompanying gestures can be part of a demonstration, or, in other words, quoted. this leads ebert & hinterwimmer (2022) to combine davidson’s (2015) proposal with the formal proposal for the semantic contribution of gestures as proposed by ebert et al. (2020) (cf. section 2.2 of the present article). by means of a rating study, they test their claims by testing dd, id, and fid speech reports in which a pronoun occurs that is co-referent with the reported speaker (first-person pronoun in dd, third-person pronoun in id and fid). a sample item (translated from german) is given below. capitalization indicates focalization. (8) a. fid: leona was extremely annoyed. again, she had paid the bill for the entire group. + self-pointing b. id: leona was extremely annoyed. she was angry that she had paid the bill for the entire group again. + self-pointing c. dd: leona was extremely annoyed. angrily, she thought: “now, i have paid the bill for the entire group again.” + self-pointing (ebert & hinterwimmer 2022, p. 345) proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 428 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ following davidson (2015) in the assumption that quotation is demonstrational and therefore also gestures can be quoted, ebert & hinterwimmer (2022) hypothesize that self-pointing aligned with the first-person pronoun should be acceptable in the dd example (8c) because it is then interpreted as a cvg quoted from the reported speaker. if maier’s (2015) theory is correct and fid in fact involves mixed quotation, the self-pointing should also be acceptable in examples like (8a). finally, for id utterances as in (8b) they hypothesize that self-pointing should not be acceptable when aligned with a third-person pronoun as id does not involve quotation and therefore the gesture cannot be interpreted as a quoted cvg. they thus predict that self-pointing should be judged significantly worse in id than in fid and dd. their results were only partly in line with their hypothesis. descriptively, the results show that id was the condition that received the worst ratings. an anova revealed, however, that this difference was only significant in the f1 and f2 analysis in the contrast between dd and id. the contrast between fid and id only reached significance in the f1 analysis, which they attribute to the surprisingly good ratings obtained for the id condition. overall, they interpret the results as follows: the good ratings for the fid condition show that it is indeed quotational, thus presenting evidence in favor of the approach proposed by maier (2015). for id, they argue, the results imply that a similar perspective-shifting strategy as in fid is available there, as well, meaning that selfpointing in id can also be interpreted as a quoted cvg. ebert & hinterwimmer’s (2022) interpretation of the findings for id thus suggests that it also is an instance of mixed quotation. it should thus receive a similar treatment as fid in the account proposed by maier (2015). this aligns with the observation that some indexical expressions can shift or obligatorily shift toward the reported speaker in id. however, the quotation mechanism should arguably be more constrained than in fid as more expressions are interpreted from the reporting speaker’s perspective in id than in fid. the study reported in this paper aims to test this hypothesis by adapting the study reported in ebert & hinterwimmer (2022). it tested whether altering the perspective prominent at the speech level in id could affect the availability of interpreting a self-pointing gesture as a cvg quoted from the matrix subject. following work suggesting that some indexicals can shift if the matrix subject’s perspective is prominent (plank 1986, anderson 2019), it is hypothesized that self-pointing gestures aligned with a third-person pronoun referring to the matrix subject are acceptable in id only if the matrix subject’s perspective is prominent on the speech level. when the reporting speaker’s perspective is prominent in the id utterance, by contrast, self-pointing should not be acceptable as it then cannot be interpreted as a cvg quoted from the matrix subject. 3. experimental study. 3.1. method. 3.1.1. participants. self-reported native speakers of german (n = 60) were recruited via prolific. all of them were naive with respect to the research question. 3.1.2. materials. sixteen videotaped experimental items were constructed. each experimental item consisted of an opening statement, followed by the target sentence rendered in id. the first sentence described how the protagonist, i.e., the matrix subject of the id sentence, felt. the subsequent id utterance explained why they felt this way. crucially, in the id sentence, either proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 429 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the matrix subject’s (cf. (9a)) or the speaker’s perspective (cf. (9b)) was made prominent (factor perspective: matrix subject vs. speaker). this was tested in a pilot study. in addition, each id sentence contained a focalized third-person pronoun which was co-referent with the matrix subject. this pronoun was either aligned with a self-pointing or a beat gesture (factor gesture: self-pointing vs. beat). thus, the study was of a 2x2 design. an example is given in (9): (9) a. matrix subject’s perspective: pia ging es erbärmlich. sie fragte sich, warum ihre beste freundin anna, diese gottverdammte saufziege, gestern abend mal wieder ihr zu viel wein nachgeschüttet hat, obwohl sie doch so wenig verträgt. + self-pointing/beat ‘pia was feeling miserable. she wondered why her best friend anna, that damned lush, had poured her too much wine again last night, even though she couldn’t handle it.’ b. speaker’s perspective: pia ging es erbärmlich. sie fragte sich, warum ihre beste freundin anna, die aber eigentlich nur die letzte pfütze aus der weinflasche loswerden wollte, gestern abend mal wieder ihr zu viel wein nachgeschüttet hat, obwohl sie doch so wenig verträgt. + self-pointing/beat ‘pia was feeling miserable. she wondered why her best friend anna, who was just trying to get rid of the last drops in the wine bottle, had poured her too much wine again last night, even though she couldn’t handle it.’ in (9a), the matrix subject’s perspective is prominent due to the appositive diese gottverdammte saufziege (‘that damned lush’). under its most salient interpretation, the appositive conveys an attitude toward anna that is most likely pia’s. it seems implausible to assume that this is the speaker’s attitude in this context. by contrast, in (9b), the appositive die aber eigentlich nur die letzte pfütze aus der weinflasche loswerden wollte (‘who was just trying to get rid of the last drops in the wine bottle’) is most plausibly interpreted as additional information about the wine-pouring scenario to which only the speaker has access, not pia. this makes the speaker’s perspective more prominent in this example. to distract participants from the experimental manipulations, the experimental items were interspersed with 30 unrelated fillers. 3.1.3. procedure. participants were first made familiar with the task by means of an introductory text. in this text, they were explicitly instructed to always pay attention to the audio and the videotape. in addition, they were informed about their data protection rights and had to give informed consent before starting to complete the questionnaire, which started with a training session consisting of two items. sosci survey (leiner 2022) was used to create the questionnaire, an online platform which can be used free of charge for academic purposes. the questionnaire was distributed via prolific using its pre-filtering function to exclude participants from the aforementioned pilot study and to ensure that only native speakers of german participate in the experiment. each list was run independently in order to prevent participants from participating multiple times. again, this was achieved by using prolific’s pre-filtering function. the items were split up according to a latin square design onto four lists. the 30 fillers and the two training items were included on every list. experimental items and fillers were randomized for each participant. their task was to rate the items on a 7-point likert scale for acceptability (1 = completely unacceptable; 7 = completely acceptable). to check whether the participants paid attention to the stimuli, they were asked attention questions about the videos (e.g., they were asked proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 430 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ what the color of the background in the videos was) at the end of the questionnaire. 3.2. predictions. based on the hypothesis that self-pointing gestures in id can be interpreted as cvgs quoted from the matrix subject when their perspective is prominent, it is assumed that self-pointing will be rated significantly higher in the matrix subject condition of the factor perspective. since beat gestures synchronize with the speech rhythm in general (mcneill 1992), they are assumed to be highly acceptable regardless of the manipulation of the factor perspective. therefore, an interaction between the factors perspective and gesture is predicted. 3.3. results. statistical analysis was done using the programming language r (r core team 2022) inside the integrated development environment rstudio (rstudio posit team 2023). for data processing and visualization, the package ‘tidyverse’ (wickham et al. 2019) was used. to test for significant effects, the results were analyzed using a cumulative link mixed effects model with the clmm() function in the r package ‘ordinal’ (christensen 2023). using effect coding, the two factors gesture and perspective and all their interactions were entered as fixed effects into the model, meaning that the intercept represents the unweighted grand mean and the fixed effects compare the factor levels to each other. the analysis script as well as the materials are available on osf: https://osf.io/h6rt5/. figure 3: mean values and standard deviations for each condition. (abbreviations: sp = selfpointing, ms = matrix subject’s perspective prominent, speaker = speaker’s perspective prominent) figure 3 shows the mean values and standard deviations. the self-pointing condition was rated equally well in both conditions of the factor perspective (matrix subject: m = 4.47, sd = 1.95; speaker: m = 4.54, sd = 1.96). the same can be observed for the beat condition for gesture (matrix subject: m = 5.04, sd = 1.68; speaker: m = 5.24, sd = 1.67). in general, it can be observed that beat gestures were overall rated better than self-pointing gestures. the output of the ordinal mixed-effects model is given in table 1. it shows a main effect for the factor gesture, confirming the tendency observed in the descriptive data shown in figure 3: beat gestures were rated significantly higher than self-pointing gestures, irrespective of the manipulation of the factor perspective. 3.4. discussion. contrary to the predicted interaction, the results only show a main effect for gesture. despite this contradiction, the results are still in line with the hypothesis: as evidenced proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 431 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ estimate std. error z value pr(> |z|) gesture -.757 .12 -6.212 5.24e-10 *** perspective .162 .118 1.375 .169 gesture:perspective -.181 .235 -.767 .443 table 1: ordinal mixed-effects model with mode and match as fixed effects and participants and items as random intercepts. formula: choiceo ∼ gesture * perspective + (1|case) + (1|item) significance codes: *** 0.001 | ** 0.01 | * 0.05 | . 0.1 by the descriptive data in figure 3, self-pointing gestures are acceptable in id. in a way, the results of the study reported here even go beyond the initial hypothesis because not only are self-pointing gestures in id acceptable when the matrix subject’s perspective is prominent on the speech level, but also when the speaker’s perspective is made prominent. as will be elaborated below, the results are in line with the assumption that—in a similar vein as in fid (ebert & hinterwimmer 2022)— a perspective shift can also take place in id utterances and consequently, self-pointing gestures aligned with third-person pronouns can be interpreted as cvgs quoted from the matrix subject. 4. general discussion and conclusion. the study reported by ebert & hinterwimmer (2022) tested whether fid utterances are an instance of mixed quotation (cf. maier 2015). under the assumption that quotation involves demonstration (clark & gerrig 1990, davidson 2015), they argue that a self-pointing gesture aligned with a third-person pronoun co-referent with the reported speaker should be acceptable as it is then interpreted as a cvg quoted from them. the results supported the mixed-quotational account of fid put forth in maier (2015). as a control condition, they also included id utterances, hypothesizing that self-pointing should not be acceptable here, since id utterances are normally argued to not involve quotation. surprisingly, however, the id condition was also rated surprisingly well, thus suggesting that some form of quotation is—or at least can be—present in id, as well. the study reported in this paper aimed to further investigate this. previous research has shown that some indexical expressions can receive a shifted interpretation if the perspective of the matrix subject is prominent in the surrounding discourse (plank 1986, anderson 2019). given the findings of ebert & hinterwimmer’s (2022), this shifted interpretation could be re-analyzed as an instance of these expressions being quoted from the matrix subject, in line with their suggestions for the self-pointing data. therefore, the present study manipulated the perspective prominent on the speech level (matrix subject’s perspective vs. speaker’s perspective), hypothesizing that selfpointing interpreted as a quoted cvg is only acceptable in id if the perspective of the matrix subject is prominent on the speech level. the results go beyond this hypothesis, as self-pointing in id was found to be acceptable regardless of the perspective prominent on the speech level. this suggests that mixed quotation is available in id. this means that id should receive a formal treatment along the lines of the one proposed for fid by maier (2015). however, as not all elements can be quoted in id, the quotation mechanism should be more constrained in id in comparison to fid, as it remains unclear for now which elements in id can be quoted and under which circumstances. the results of a study reported elsewhere (walter 2024) suggest that face emoji, which normally receive an author-oriented interpretation by default (grosz et al. 2023), can also receive proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 432 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ a shifted interpretation in id. i argued that they should be analyzed as quoted facial expressions from the matrix subject. for emoji to shift, however, the matrix subject’s perspective has to be prominent, highlighting the importance for future research to further investigate the influence perspective prominence in id has on the availability of shifted interpretations and thus the availability of mixed quotation in id. references anand, pranav & andrew nevins. 2004. shifty operators in changing contexts. in robert b. young (ed.), proceedings of semantics and linguistic theory (salt) 14, 20–37. ithaca, ny: clc publications. 10.3765/salt.v14i0.2913. anderson, carolyn j. 2019. tomorrow isn’t always a day away. in m. teresa espinal, elena castroviejo, manuel leonetti, louise mcnally & cristina real-puigdollers (eds.), proceedings of sinn und bedeutung 23, 37–56. barcelona, spain: universitat autònoma de barcelona. beattie, geoffrey & heather shovelton. 2002. an experimental investigation of some properties of individual iconic gestures that mediate their communicative power. british journal of psychology 93(2). 179–192. 10.1075/gest.1.2.03bea. christensen, rune h. b. 2023. ordinal—regression models for ordinal data. https://cran. r-project.org/package=ordinal. r package version 2023.12-4. clark, herbert h. & richard j. gerrig. 1990. quotations as demonstrations. language 66(4). 764–805. 10.2307/414729. davidson, donald. 1979. quotation. theory and decision 11(1). 27. 10.1007/bf00126690. davidson, kathryn. 2015. quotation, demonstration, and iconicity. linguistics and philosophy 38(6). 477–520. 10.1007/s10988-015-9180-1. ebert, christian, cornelia ebert & robin hörnig. 2020. demonstratives as dimension shifters. in proceedings of sinn und bedeutung 24, 161–178. osnabrück: university of osnabrück. ebert, cornelia. 2024. semantics of gesture. annual review of linguistics 10(1). 169–189. ebert, cornelia & stefan hinterwimmer. 2022. free indirect discourse meets character viewpoint gestures. in sam featherston, robin hörnig, andreas konietzko & sophie von wietersheim (eds.), proceedings of linguistic evidence 2020: linguistic theory enriched by experimental data, 333–349. tübingen, germany: university of tübingen. eckardt, regine. 2014. the semantics of free indirect disocurse: how texts allow to mind-read and eavesdrop. leiden, the netherlands: brill. 10.1163/9789004266735. farkas, donka f. & kim b. bruce. 2010. on reacting to assertions and polar questions. journal of semantics 27(1). 81–118. 10.1093/jos/ffp010. grosz, patrick g., gabriel greenberg, christian de leon & elsi kaiser. 2023. a semantics of face emoji in discourse. linguistics and philosophy 46(4). 905–957. harris, jesse a. 2012. processing perspectives. amherst, ma: university of massachusetts amherst dissertation. hinterwimmer, stefan. 2017. two kinds of perspective taking in narrative texts. in dan burgdorf, jacob collard, sireemas maspong & brynhildur stefánsdóttir (eds.), proccedings of semantics and linguistic theory (salt) 27, 282–301. university park, md: university of maryland. 10.3765/salt.v27i0.4153. proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 433 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ kendon, adam. 2004. gesture: visible action as utterance. cambridge, uk: cambridge university press. 10.1017/cbo9780511807572. leiner, daniel j. 2022. sosci survey (version 3.3.14a). available at https://www.soscisurvey.de. maier, emar. 2015. quotation and unquotation in free indirect discourse. mind & language 30(3). 345–373. 10.1111/mila.12083. mcneill, david. 1992. hand and mind: what gestures reveal about thought. chicago, il: university of chicago press. parrill, fey. 2010. viewpoint in speech-gesture integration: linguistic structure, discourse structure, and event structure. language and cognitive processes 25(5). 650–668. 10.1080/01690960903424248. plank, frans. 1986. über den personenwechsel und den anderer deiktischer kategorien in indirekter rede. zeitschrift für germanistische linguistik 14(3). 284–308. 10.1515/zfgl.1986.14.3.284. potts, christopher. 2005. the logic of conventional implicatures. oxford, uk: oxford university press. 10.1093/acprof:oso/9780199273829.001.0001. quine, willard v. 1940. mathematical logic, vol. 4. cambridge, ma: haravrd university press. r core team. 2022. r: a language and environment for statistical computing. r foundation for statistical computing vienna, austria. https://www.r-project.org/. rstudio posit team. 2023. rstudio: integrated development environment for r. posit software, pbc boston, ma. http://www.posit.co/. de ruiter, jan p. 2000. the production of gesture and speech. in david mcneill (ed.), language and gesture, cambridge, uk: cambridge university press. 10.1017/cbo9780511620850.018. schlenker, philippe. 2004. context of thought and context of utterance: a note on free indirect discourse and the historical present. mind & language 19(3). 279–304. 10.1111/j.14680017.2004.00259.x. tonhauser, judith. 2012. diagnosing (not-)at-issue content. in elizabeth bogal-allbritten (ed.), proceedings of semantics of under-represented languages of the americas (sula) 6, 239– 254. walter, sebastian. 2024. can face emoji receive a shifted interpretation in indirect discourse? poster presented at architechtures and mechanisms of language processing (amlap) 30, edinburgh, uk. wickham, hadley, mara averick, jennifer bryan, winston chang, lucy d’agostino mcgowan, romain françois, garrett grolemund, alex hayes, lionel henry, jim hester, max kuhn, thomas lin pedersen, evan miller, stephan milton bache, kirill müller, jeroen ooms, david robinson, dana paige seidel, vitalie spinu, kohske takahashi, davis vaughan, claus wilke, kara woo & hiroaki yutani. 2019. welcome to the tidyverse. journal of open source software 4(43). 1686. 10.21105/joss.01686. proceedings of elm 3: 423-434, 2025 sebastian walter: indirect discourse as mixed quotation? an experimental investigation. 434 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ experientiality markers in memory reports: a semantics-pragmatics puzzle emil eva rosina & kristina liefke* abstract. some recent work in semantics and the philosophy of language suggests that the way we report events reflects whether we have personally experienced or witnessed these events (i.e. through linguistic elements dubbed ‘experientiality markers’). this paper provides experimental support for one such marker: german non-manner uses of wie [‘how’]. we argue that when they are embedded under the memory predicates noch wissen [‘still know’] and sich erinnern [‘refl-remind’], free relative wiecomplements mark the remembering of a personally experienced event. we support this claim through a series of online studies based on scale judgements. the results of our main study raise questions about the semantics-pragmatics interface of the experientiality marking property of wie, and about the robustness of experientiality markers in general. a series of complementary studies address these questions. keywords. experiential remembering; memory predicates; attitude reports; knowledge; evidentiality; study formats; propositional attitudes; pragmatic competitions 1. introduction. german memory predicates can combine with a declarative dass-[‘that’-]clause (1-b) and with an eventive-wie [‘how’] free relative (1-a).1 this holds both for the reflexive predicate sich erinnern (lit. ‘oneself remind’) and for the complex noch wissen (lit. ‘still know’). (1) a. ich i {erinnere remind mich/ myself/ weiß know noch}, still wie how oma grandma im in-the meer sea geschwommen swim ist. is ‘i remember grandma swimming in the sea.’ b. ich i {erinnere remind mich/ myself/ weiß know noch}, still dass that oma grandma im in-the meer sea geschwommen swim ist. is ‘i remember that grandma was swimming in the sea.’ (german) interestingly, dass-clauses and wie-free-relatives can be coordinated under either of the above predicates. following familiar ambiguity tests (cf. sadock & zwicky 1975), we thus assume a uniform semantics for noch wissen in (1-a) and (1-b) and another uniform semantics for sich erinnern in (1-a) and (1-b), such that these sentences use the same semantic entry for the matrix verb and hence form minimal pairs in our studies. in rosina & liefke (2024a), we give such a unified (and fully compositional) semantics of noch wissen as retained knowledge – of an informationally rich *we thank yonca klisch, jonas koopmann and robert kurth for their help with the study design. we thank markus werning, nina haslinger, justin d’ambrosio, alex wiegmann, sebastian walter, deniz özyıldız, and jan köpping, and the audiences and reviewers of elm 3, sinfonija, wccfl, dgfs, the frankfurt semantics colloquium, and the mecore closing workshop for their helpful comments. authors: emil eva rosina, ruhr university bochum (emil.rosina@ruhr-uni-bochum.de) & kristina liefke, ruhr university bochum (kristina.liefke@ruhr-unibochum.de). the research for this paper is supported by the german research foundation dfg as part of the research unit for 2812 constructing scenarios of the past (grant no. 397530566). 1german only admits two of the three uses of how that have been attested for english: next to the familiar mannerwie (‘the way in which’), it only allows for eventive uses of non-manner wie (the sole focus of the present paper; cf. umbach et al. 2022, liefke 2023). factive uses of how (‘they told me how the tooth-fairy exists’, cf. legate 2010) are not attested for german. proceedings of elm 3: 319-331, 2025 c©2025 emil eva rosina and kristina liefke published by the lsa with permission of the author(s) under a cc by license. 319 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (wie) or poor (dass) proposition. our semantics suggests (but does not explicity claim) an extension to (stative uses of)2 sich erinnern, english remember and other memory predicates, such that still+know is the core of remembering and noch wissen is just the most transparent spell-out. the present paper presents a series of online behavioural studies. three of these studies (the ones presented in sect. 2) vary with respect to the (german and english) memory predicates used in the test sentences. their very similar results suggest that these predicates may indeed share a semantic core that interacts with different complements in a uniform way across the concrete spellout of these predicates. concerning the truthand utterance-conditions of sentences like (1-a) and (1-b), a satisfactory semantic account has to position itself relative to philosophical, psychological, and neuroscientific work on memory. this work (going back to tulving 1972) commonly distinguishes experiential (‘episodic’) remembering (i.e. recall of a personally experienced event) from fact-only (‘semantic’) remembering, i.e. recall of general facts, often based on indirect evidence or testimony. in our experiments, we introduce the siblings red and blue to personify these kinds of experience. in particular, we tell our participants the following about them in the case of the swimming scene that relates to the sentences in (1): (2) a. red spent the summer two years ago with grandma and saw her swimming in the sea. b. blue spent that summer abroad and was told about grandma’s swimming much later. based on our semantics in rosina & liefke (2024a) and in line with literature on non-manner uses of how (liefke 2023, umbach et al. 2022), we expect that (1-a) unambiguously reports experiential memory (red, (2-a)) while (1-b) is expected to report both fact-only (blue, (2-b)) and experiential memory (cf. fig. 1). by confirming this, we provide the first empirical evidence for experientiality markers in memory reports. the idea that how we report events reflects whether we have personally experienced or witnessed these events (i.e. through particular linguistic elements dubbed ‘experientiality markers’) is found in bernecker (2010) a.o. beyond its contribution to semantics, the empirical identification of experientiality markers in memory reports might be taken to provide further evidence for the two philosophically distinct types of remembering (i.e. experiential and fact-only remembering, going back to tulving 1972), given some bridging principles. additionally, being able to pinpoint specific experientiality markers like german wie is particularly useful for analyzing production data from psychological memory studies that did not themselves create the remembered (and subsequently reported) experience. an important feature of rosina & liefke (2024a) is that, in this account, experientiality is not directly encoded in the semantics. instead, this account attributes low acceptance rates of blue’s (2-b) ‘remember how’-equivalents to blue’s lack of good evidence for the informationally rich proposition ‘how p’ (roughly: ‘that the things were such-and-such when p in @’; see rosina & liefke 2024a for the interplay of informational richness and evidence-based knowledge). the intuition that one must have personally experienced an event in order to truthfully self-attribute ‘remembering how’ is an indirect effect of our world knowledge that direct experience is usually the best kind of evidence. low ratings for blue+how-sentences by themselves do not distinguish between this account, an account that views direct (as opposed to just good) evidence as a require2our semantics cannot, in its current form, yet capture eventive ‘is remembering right now’-uses that some memory predicates have in addition. proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 320 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ment, and an account that writes the requirement of personal experience directly into the semantics (stephenson 2010, liefke & werning 2024) – because blue lacks all of these. to our knowledge, we are the only ones choosing the first option (good evidence), while the second (direct evidence) is prominent in related literature on perception (see e.g. davis & landau 2021 on see that vs. see -ing), and the third (personal experience) in previous philosophical and semantic accounts of experiential memory (stephenson 2010, liefke & werning 2024). our three studies targeting the semantics-pragmatics interface of memory reports, including this issue, are presented in sect. 3. for a more detailed discussion of possible pragmatic effects and the inter-disciplinary significance of memory reports, see rosina (2024). 2. attesting experientiality markers. 2.1. german main study. in this section, we present our biggest study as a case of our general experimental paradigm. in the later sections, we introduce a series of complementary studies that are all variations of this main study in different respects. if not indicated otherwise, the experimental design is as described for the main study. more specifically, all studies except the qud study (see sect. 3.2) share the same basic background story of a family gathering (introduced below). with the exception of the speaker-id study (see sect. 3.1), all studies share the rating format, and all except the intermediate evidence study (see sect. 3.3) share the same set of characters. alongside our target characters red and blue (see ex. (2) in sect. 1), we introduce their cousin pinkie for controls and tell the participants that pinkie does not have any evidence concerning the events depicted in our target vignettes. the three teenage cousins represent different types of experience/evidence with respect to different past events of some mishap involving their grandmother. fig. 1 introduces the background story and the characters’ viewpoints on one example event. figure 1: composition of screenshots and text from the english version of the experiment against this background (presented in german), our main experiment targets the dass/wiecontrast in german memory reports with the verb sich erinnern (‘refl-remember’). both the character uttering a sentence (red, blue) and the complementizer (wie, dass) are manipulated variables. by combining the values of these variables, we obtain four target items for each scene: proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 321 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (3) a. red sagt: ich erinnere mich, dass oma überfallen wurde. b. blue sagt: ich erinnere mich, dass oma überfallen wurde. c. red sagt: ich erinnere mich, wie oma überfallen wurde. d. blue blue sagt: says: ich i erinnere remind mich, refl wie how oma granny überfallen robbed wurde. was ‘{red/blue} says: i remember {that/how} granny was robbed.’ (german) our main experiment is a qualtrics online rating study that asks participants to judge sentences of this form against the background of a given scenario, consisting of ‘what happened’ and the speaker’s mnemonic perspective on it, as exemplified in fig.1.3 participants were asked to provide ratings on “der grün hinterlegte satz, von [speaker] gesagt, beschreibt die situation ...” – ‘the sentence marked in green, uttered by [speaker], describes the situation ...’ – on a scale from 1 (gar nicht richtig, ‘not correctly at all’) to 7 (völlig richtig, ‘absolutely correctly’).4 based on our background assumptions and literature-informed expectations (see sect. 1), we formulated two hypotheses before preregistering our study (rosina & liefke 2024b) and then collecting results. both hypotheses together would show that wie in ‘sich erinnern’-reports is an experientiality marker in the sense that it disambiguates for experiential memory in contrast to ‘sich erinnern, dass’. (4) a. hypothesis i: higher ratings for the red+wie than for the blue+wie condition!*** b. hypothesis ii: higher ratings for blue+dass than for blue+wie !*** we recruited participants via prolific and gathered data from 60 german mono-lingually raised native speakers aged 18–65, six of whom we excluded from analysis based on a control performance of ≤ 75%. the combination of our two manipulated variables (speaker and complementizer) resulted in four conditions, exemplified by (3). alongside the robbery scenario, we used three more target scenarios: grandma swimming and almost drowning, falling out of a canoe, and burning red’s birthday cake. hence, our study consisted of 16 target items, augmented with 16 control items.5 we tested within-subjects in order to facilitate a-posteriori reasoning. (e.g.: are there two kinds of qud-accommodators?) as a result, each participant provided four judgements per condition, resulting in 216 data points per condition. hypothesis i was clearly confirmed with an extremely strong contrast (see fig. 2; see tab. 1 for interaction). hypothesis ii was also confirmed, but with a weaker contrast due to the lower-than-expected rating of blue+dass.6 we consider these results evidence for wie as an experientiality-marker (in the semantics/ 3a mock version of the german main study can be accessed via https://bochumpsych.eu.qualtrics.com/jfe/form/sv ezvoyefvfcy04iu. 4we decided for these instructions and for this naming of the endpoints of our scale after comparing all similar experiments in the proceedings of elm1 and elm2. terminology involving accuracy or ‘good fits’ tends to be more sensitive to pragmatic and socio-linguistic effects, while asking participants to grade truth has a non-trivial ontological flavor. cf. zhu & ahn (2023) for the effect of instructions. 5false control sentences attribute pinkie remembering of the events involving grandma. true controls attribute red perception of these events, or are unrelated true statements about pinkie like ‘i am thinking about cats’ (as depicted by the pinkie-picture, see fig. 1). the target and control items are randomized within 8 blocks in order to prevent certain orders, and so is the order of the blocks. 6all analyses presented in this paper were done using cumulative link mixed effect models fitted with the laplace approximation, with participant and scene as random intercepts. (for motivation of the choice, see liddell & kruschke 2018); significance codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 322 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 2: ratings by condition and quartiles, main study estimate std. error z-value pr(> |z|) speaker:marker 0.9755 0.3014 3.237 0.00121 ** wie:red 4.2446 0.2557 16.597 <2e-16 *** blue:wie -0.8313 0.1741 -4.775 1.8e-06 *** red:wie 0.1442 0.2455 0.588 0.557 table 1: interactions of speaker and marker, main study pragmatics neutral sense, see sect. 1) in german ‘sich erinnern’-reports, in line with our semantics from (rosina & liefke 2024a), but not exclusively so. importantly, the main experiment by itself does not provide any conclusive evidence on how lexically specific the effect is (either to the predicate sich erinnern or to wie-clauses). the experiments presented in sect. 2.2 and 2.3 will suggest that evidentiality marking is indeed a cross-structural, cross-linguistic phenomenon. that blue+dass scored much lower than red+wie in our main study is a surprise: since blue+wie has even lower ratings than blue+dass (so there is a huge main effect of blue), there are participants who do not grant blue any kind of remembering even though she has reliable indirect evidence. a look at the individual participants’ widely distributed ratings of blue+dass in fig. 2 suggests a divide: while one group of participants is in line with our semantic-pragmatic explanation above, there is a second group whose members seem to have stricter conditions on memory. for members of this group, only experiential remembering is ‘real’ remembering in general – in the rating format of our main experiment, that is – see sect. 3.1 for further discussion of this possible explanation and of pragmatic effects that might be intervening here. for now, note that we do not know yet what it is about red that makes them a good speaker of wie-sentences. it could be direct witnessing or particularly good evidence (as predicted by our semantics in rosina & liefke 2024a) at this point. we will come back to this issue in sect. 3.3. finally, note that the results leave room, in principle, to reason that the low blue+wie ratings could instead be due to pragmatic competition with ‘sich erinnern, dass’ which could be preferred for independent reasons for blue. however, the most straightforward spell-out of such a pragmatic account would have to view dass as an indirectness marker. the high ratings of red+dass speak against such an account. our speaker-id study (see sect. 3.1) will shed some more light on these considerations. 2.2. english study. our study of experientiality marking through eventive uses of how has, until now, focused exclusively on german. to see whether this marking is cross-linguistically more robust (at least in a minimal sense), we have conducted a follow-up study that replicates the main proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 323 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ study for (american) english. the resulting english rating study with 27 participants (after exclusions) is a first hint that it might be quite robust.7 we did not preregister this study. it is mostly an english translation of half of the main experiment, but with the memory predicate remember and the hypothesized marker gerundive -ing small clauses (‘gsc’) instead of wie/how-clauses.8 these gerundive -ing-constructions are generally considered the prototypical way of reporting experiential remembering in english (see stephenson 2010, bernecker 2010). bernecker (2010)’s claim that experientiality of rememberings is grammatically encoded refers to english gsc specifically. importantly, bernecker (2010) also holds the inverse of this claim, viz. that that-clauses are nonexperientiality/indirectness markers. our english study (i) provides evidence against this latter claim, and (ii) will show english gsc-constructions to have more-or-less the same effect as eventive wie-[‘how’-] constructions in german. more complex than in the case of german wie/dassclauses, that-clauses can differ from gsc in two ways besides the presence/absence of that: the that-clause can feature a past progressive or past simple verb form. the past progressive makes use of the -ing form like the gsc does, so it seems the better candidate for a minimal pair. on the other hand, some events may not be naturally reported with progressive aspect (leading to pragmatic disturbance),9 and the -ing form itself may turn out to be an experientiality marker, as opposed to the small clause character of gsc. for these reasons, we decided to include two versions of that-clauses for each gsc, leading to ‘minimal triplets’ (5) and six conditions in total.10 (5) a. red says: i remember that grandma got robbed. b. blue says: i remember that grandma got robbed. c. red says: i remember grandma getting robbed. d. blue says: i remember grandma getting robbed. e. red says: i remember that grandma was getting robbed. f. blue says: i remember that grandma was getting robbed. the participants judged these sentences against only two of the target scenes from the main study (robbery and swimming), such that they answered 12 target items and 12 controls, resulting in 54 analysed data points per condition. (this explains the lower significance of the effects relative to the main study.) the english instructions read as follows: “the sentence highlighted in green, said by [speaker], describes the situation... 1 (not correctly at all) ... 7 (completely correctly)”. we analyzed the pairings gsc/that-simple and gsc/that-ing separately and confirmed hypotheses ib and iib in (6) when the that-clause uses past simple as in (5-a-b). (6) a. hypothesis ib: higher ratings for red+gsc than for blue+gsc !*** 7since the kind of non-manner uses of how that are relevant here seem more common in american english (see liefke 2023), we conducted this study with monolingual-native english speakers currently living in the us. the only other screener on prolific was age (18–65). 8a mock version of the english study can be accessed via https://bochumpsych.eu.qualtrics.com/jfe/form/sv afnnktpbb9wvmva. 9we thank justin d’ambrosio, p.c., for pointing us to this. 10we are aware of the in-principle syntactic ambiguity of the -ing-constructions in (5-c-d) between a gscconstituent and a dp+adjunct (equivalent to ‘i remember grandma, as she was getting robbed’). the syntax of the construction we dub ‘gsc’ is irrelevant for its status as an experientiality marker for now, but will have to be considered for a future compositional account. proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 324 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ b. hypothesis iib: higher ratings for blue+that-simple than for blue+gsc !* the results of this second study show a significant interaction between marker and speaker. the effects are not quite significant for ‘. . . that grandma was getting robbed’ (5-e-f), suggesting that -ing itself contributes to experientiality/evidentiality. comparing the results in fig. 3 with the results of the main study in fig. 2, the close parallel is evident. figure 3: ratings by condition and quartiles, english study first, note that red+that-simple has very high ratings, and there is no significant contrast with red+gsc. (this part also holds for red+that-ing.) this is the first clear empirical evidence against bernecker (2010)’s claim that that-clauses mark indirectness of evidence/experience. the absence of any significant preference for red+gsc/wie over red+that/dass in the two studies is also interesting in the light of possible pragmatic competition: if both options are open for the direct experiencer red, semantically, one might expect pragmatically decreased ratings for red+that/dass due to ‘maximize precision’, because blue can use that/dass as well, but not gsc/wie (cf. grice 1975). we observe no such effect in any of the rating studies,11 and will return to this point when discussing the speaker-id study in sect. 3.1. a link with the blue+that/dass puzzle suggests itself: if for whatever reason blue cannot acceptably utter any memory report, the choice of the complement does not in fact maximize precision from red’s perspective. coming back to the general picture, the results of the english study – including the puzzle on the lower-than-expected ratings for blue+dass (now blue+that-simple) – closely resemble the german results from the main study. this suggests an at least minimal robustness of experientiality marking across languages (german and english), memory predicates (sich erinnern and remember), and complement structures marking experientiality (wie/how-clauses and gsc). 2.3. german ‘noch wissen’ study. this case becomes even stronger when we consider the results of a third study that only differs from the german main study in the following respects: (i) most importantly, the matrix predicate is noch wissen [‘still know’] instead of sich erinnern [‘refl-remember’]. investigating the parallel or different distribution of these two predicates is motivated by the considerations discussed in sect. 1 and our rosina & liefke (2024a), suggesting that noch+wissen could be the core of memory predicates in general. if noch wissen shows the same empirical pattern as sich erinnern, this motivates a closely related semantics, minimally concerning the selectional flexibility and the relation to wie-complements and experientiality. (ii) this was a smaller-scale study, with 37 participants after exclusions and no preregistration. 11this can also be attributed to our instructive formulations which were – it seems, successfully – aimed at truthconditional effects as far as possible in an acceptability experiment, see fn. 4. proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 325 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (iii) like for the english study, we used only half of the scenes from the main study to minimize participation time for participants. this lead to 8 target items (2 per condition) plus 8 controls and 74 data points per condition. our results confirm both hypotheses in (7) and significant interaction of marker and speaker. (7) a. hypothesis ic: higher ratings for red+wie than for blue+wie !*** b. hypothesis iic: higher ratings for blue+dass than for blue+wie !*** besides providing tentative support for our idea that our semantics for noch wissen in rosina & liefke (2024a) may be generalized, the combined results of the three experiments presented so far target the relation between language, evidence, and cognition at its core. a first, modest conclusion is that an account of ambiguity or polysemy of individual memory verbs has become extremely unattractive in the light of the (weakly, so far) cross-linguistic and cross-predicate (also within german) picture, and we should aim at finding a core, unified semantics of memory reports. the idea would be to give remember-equivalents a unified semantics that gives a different output in terms of use-conditions for wieand gsc-complements on the one hand and dass/thatcomplements on the other hand. our rosina & liefke (2024a) achieves this in a compositional manner for the case of wie/dass. to extend our coverage to the english results, gsc would have to be analyzed as informationally rich propositions like wie-complements. the above considerations suggest a more far-reaching – but also more speculative – parallel with other attitude predicates and with perception predicates. the first part of this parallel lies in the presence of a clear relation between noch wissen and wissen [‘know’], such that (even experiential) remembering is a kind of knowledge (cf. hörl 2022). the effect of informational richness on evidence is present also in the case of present-time imaginative knowledge. the speaker must be constructing an informationally rich scenario in their mind, and have very good (most likely, experiential) evidence for ‘the way things are’ in the neighbors’ house at the time of uttering (8-b). (8) a. having watched the neighbors snap at each other and storm into their house: b. ich i weiß know (genau/schon/ja/#noch), (exactly/part/part/still) wie how die these jetzt now (wieder) (again) aufeinander at-each-other losgehen. attack ‘i know (exactly) how these two are fighting (again) right now.’ r&l (2024a) ex. (14) davis & landau (2021) investigate a similar effect for gsc under perception verbs. perhaps this general interaction of informational richness, experience, and evidence can be described in a way uniform across these very different predicates. for an account in the spirit of our rosina & liefke (2024a), they would all have to share an evidential core which interacts pragmatically with informational richness (which is, again, encoded by different structures) in a way that increases the required quality of evidence the more informationally rich the object of this evidence is. this constitutes further support that cognitive concepts of evidentiality ‘precede’ its natural language realizations (ünal & papafragou 2018), but in a way that does not require any clear categorisation of direct vs. indirect evidence, since it is only about good vs. bad evidence. we will return to the issue of the gradient nature of evidence in sect. 3.3, but leave the general discussion of how broad the phenomenon is open for future work. proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 326 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3. experiments on pragmatic effects. we now return to another route of open issues from sect. 1 and the main experiment discussed in sect. 2.1 that have only been strengthened by the english study and the noch wissen study. remember that we could not conclusively distinguish between truth-conditional and use-conditional effects. in particular, red+that/dass was rated significantly higher than blue+that/dass, due to an extremely significant main effect of speaker. for some participants, the indirect experiencer blue seems to be granted no remembering at all. three complementary studies address these issues. 3.1. speaker-id study. the set-up and the phrasing of the five experiments were aimed at truth-conditional semantics (for the influence of instructive formulations on results, see zhu & ahn 2023). in an attempt to control for pragmatic competition, we ran a smaller, non-preregistered experiment in another format with 29 german-native participants after exclusions and four target plus four control items.12 since this format only has two conditions, this amounts to 58 data points per condition. in this speaker-identification format (inspired by davis & landau 2021), participants choose “wer sagt [sentence]? – red blue pinkie” (‘who says [sentence]?’ – red, blue or the control character pinkie)13 for each of the sentences in (9) presented in a given scene (cf. fig.1). (9) a. ich weiß noch, dass oma überfallen wurde. b. ich i weiß know noch, still wie how oma granny überfallen robbed wurde. was ‘i remember {that/how} granny was robbed.’ c. ich weiß noch, dass oma im meer geschwommen ist. d. ich i weiß know noch, still wie how oma granny im in-the meer sea geschwommen swim ist. is ‘i remember {that/how} granny was swimming in the sea.’ (german) note that only the complementizer varies between (9-a/c) and (9-b/d), since the character is now selected instead of given. this leads to a 2x2 setup with the speaker as the dependent variable and the complementizer the only manipulated variable. we forced participants to decide for exactly one character, instructing them to choose the one who is more likely to have uttered the sentence, if more than one or none of them could have said it. the results confirm both hypotheses in (10): (10) a. hypothesis i: red is selected more often given the wie-condition !84% b. hypothesis ii: blue is selected more often given the dass-condition !64% like the results of the rating studies, these provide evidence for some version of the claim that german wie-complements mark experientiality. two more things are noteworthy about the speakerid results: first, the effect of blue>red for the dass-condition is weaker than the effect of red>blue for the wie-condition. since truth-conditions show stronger effects in experiments than competitions do, this supports the status of wie as a semantic marker (of direct experience or good evidence, see sect. 3.3) in opposition to dass, which is not an indirectness marker, semanti12a mock version of the german speaker-id study can be accessed via https://bochumpsych.eu.qualtrics.com/jfe/form/sv 5j8cj1jehaak5m2. 13no participant selected pinkie for any of our target items, which facilitates the 2x2-analysis. proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 327 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ cally, but is assigned more often to blue because of pragmatic competition: blue does not have any other way to express her remembering, because the wie-sentence is excluded qua truth-conditions; red could have used the wie-sentence instead to maximize precision. while this explanation is intuitively plausible when considered in isolation, it is in tension with what we observed for the rating studies. even the confirmation of hypothesis ii is unexpected relative to the results of all five rating studies: the preference for blue>red in the dass-condition clashes with significantly higher ratings for red+dass than for blue+dass. if participants are presented with the german equivalent of the sentence ‘i remember that grandma was swimming in the sea’, they choose blue as a speaker. however, if we ask participants about this very sentence uttered by blue, they rate it lower than a version uttered by the direct experiencer red. remember that after the main experiment, we hypothesised that many participants have very strict conditions of remembering, which exclude all cases but experiential remembering, regardless of the complementizer. however, if experiential remembering was just always ‘the real’ memory, we wouldn’t expect any preference for blue in the dass-condition in the speaker-id experiment, but rather a similarly strong preference for red in both conditions. directly after analyzing the speaker-id results, we suspected that the forced choice design gives rise to the pragmatic competition we intended while our judgement scale design is more sensitive to the accommodation of different questions under discussion (qud; simplifying: the purpose of the conversation). this could have been a weakness of the general design of our experimental paradigm: that the grandchildren are said to exchange stories of the old time (see fig. 1) might lead some people to accommodate a qud like ‘who was there when that happened?’. these considerations motivated the study presented in the next section. 3.2. qud-manipulated study. in attempt to test for a possible effect of the qud on the results of the rating studies, we contrasted the main experiment with a version that introduces a fact-based qud like ‘who knows the most facts about grandma?’ (preregistered: rosina & liefke 2024d). in order to enable this comparison, the target items themselves were exactly the same as in the main experiment, with sich erinnern as the matrix predicate and all 16 target and 16 control items. we made only the following changes: (i) we changed the background story such that it keeps the characters, their experiences, and the family constellation constant, but now introduces a quiz context instead of an informal conversation at a family gathering. the goal of the quiz game is to utter as many true facts about grandma as possible. the reasoning behind this was that the original background story could have made the qud ‘who had which direct experiences with grandma?’ salient, leading to the speaker main effect and the impression that blue does not have good memory at all, compared to red, even in the fact-only dass case. while the target items remained unchanged in every other respect, a reminder that the teenagers are participating in this kind of quiz was displayed with all items. (ii) we replaced all of the original control items (see fn. 5 for the nature of the original ones), introducing completely new scenes and pictures. the new controls were aimed to support the factbased qud by having as little connection to the teenagers’ personal past as possible. examples of this are sentences about grandma’s tattoo, previous jobs, or the color of her motorcycle. (iii) we recruited only 40 participants and excluded 2 of them, so the conditions qud-fact and qud-original are unevenly distributed in the pooled analysis. proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 328 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the original hypotheses i and ii were reproduced with the new background story (as id and iid below) without any noticeable differences between the two studies. to our surprise, hypothesis iii, which had assumed an effect of the qud-manipulation, was falsified. (11) a. hypothesis id: higher ratings for red+wie than for blue+wie !*** b. hypothesis iid: higher ratings for blue+dass than for blue+wie !*** c. hypothesis iii: significant reduction of the main effect red>blue in the qud-fact condition of a pooled analysis with the results of the main study % the fact that the qud does not effect any difference between the formats has two consequences for our theory formation: first, having excluded the first candidate, we have to keep looking for the reason for the difference between the scale judgement format and the speaker-id format with respect to the case of blue+dass. we suspect that the two formats support/block different pragmatic competitions, an observation that may have far-reaching consequences for methodology at the semantics-pragmatics interface in general, and that is discussed a bit more in rosina (2024). second, we are supported in our original experimental design and conclude that the background story was not a disturbing factor in any sense in the other experiments. 3.3. intermediate evidence study. concluding the series of experiments, we designed a version of the main study (preregistered: rosina & liefke 2024c) with 36 german participants after exclusions. this ‘intermediate evidence’-study aims to distinguish between accounts of experientiality markers that encode direct evidence or experience directly in the semantics (stephenson 2010, liefke & werning 2024) and our rosina & liefke (2024a) account that locates experientiality at the level of pragmatics and claims that eventive wie [‘how’] in isolation is – semantically speaking – only a marker of informational richness. in rosina & liefke (2024a), we show that when eventive wie is embedded under predicates of knowledge and remembering, the combined semantics marks good enough evidence (regardless of the kind of evidence). like in the main study, the memory predicate in this study is sich erinnern [‘refl-remember’], but we used only two target scenes (here: grandma swimming in the sea and burning a cake), like in the ‘noch wissen’ study. the most important change we made for this study is that we extended our cast of characters (see ex. (2) and fig. 1) by one additional character, red and blue’s cousin goldie. the idea behind goldie is that she has evidence that lies between red’s and blue’s in terms of ‘quality’/reliability, but no direct experience. specifically, to create this kind of evidence we told participants that goldie always missed the things that happened to grandma by a few minutes, but saw the immediate result (e.g. grandma’s broken leg or smoke in the kitchen). adding goldie as a value of speaker, and maintaining the values wie and dass for the marker variable as in all german rating studies results in six conditions with 72 data points each. based on our results from the previous studies, we reasoned that ratings between red’s and blue’s for goldie in the wiecondition would confirm the gradability of the concept licensing experientiality/evidence marking. this is exactly what we found. there were highly significant main effects of red>goldie>blue, and the dass/wie-contrast shrinks ‘in the red direction’, see our hypotheses in (12) and fig. 4. (12) a. hypothesis ie: higher ratings for red+wie than for blue+wie !*** b. hypothesis if: higher ratings for red+wie than for goldie+wie !*** proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 329 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ c. hypothesis ig: higher ratings for goldie+wie than for blue+wie !*** d. hypothesis iie: higher ratings for blue+dass than for blue+wie !*** e. hypothesis iif: higher ratings for goldie+dass than for goldie+wie % pr(> |z|) = 0.114 f. hypothesis iig: the effect confirming iie is bigger than the effect confirming iif ! figure 4: ratings by condition and quartiles, ‘intermediate evidence’-study these results provide support for our semantics in rosina & liefke (2024a) and against accounts that rely on direct experience or a specific kind of evidence to license experientiality markers in memory reports. however, proponents of such accounts could of course re-conceptualize their core concepts as gradient (e.g. based on a scale of directness of experience). we view our rosina & liefke (2024a) account as more straightforward, because it requires no mapping from kinds of evidence to a scale in the semantics. the relationship between evidence and informational richness on our account is built on one principle: the more informationally rich, the harder to be evidentially supported. (for more discussion of this, see rosina 2024.) 4. conclusion. our studies confirm the common assumption that the way we report events reflects whether we have personally experienced or witnessed these events. we provide experimental support for two such markers of experientiality: german eventive wie [‘how’] and english gerundive -ing small clauses. our studies further show that the effect of experientiality marking is not specific to any particular memory predicate, since the results for such marking under sich erinnern, noch wissen and remember are very similar. relating these results to compositional accounts of memory reports, our ‘noch wissen’-study provides tentative evidence for still+know as the core of memory predicates, as suggested in rosina & liefke (2024a). the results of our ‘intermediate evidence’-study support our idea that only informational richness and (quality of) evidence are semantically encoded, and the common requirement of direct experience is only a pragmatic effect. interestingly, even dass/that-sentences uttered by our indirect evidence character blue receive relatively low ratings, compared to the same sentence uttered by our direct experiencer red. far from marking non-experientiality (bernecker 2010), our rating studies suggest that the status of ‘remembering that’ as reporting (also) fact-only remembering is questionable. this is in tension with the results of our speaker-id study, where participants chose the indirect experiencer to be more likely to have uttered the ‘remember that’-equivalent. we suspect that the two formats give rise to different kinds of pragmatic competition and leave the formalisation of this effect to future work. proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 330 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ references bernecker, sven. 2010. memory: a philosophical study. oxford: oxford university press. davis, e. emory & barbara landau. 2021. seeing vs. seeing that: children’s understanding of direct perception and inference reports. in proceedings of elm 1, 125–135. grice, h. paul. 1975. logic and conversation. in peter cole & jerry morgan (eds.), speech acts [syntax and semantics 3], 41–58. academic press. hörl, christoph. 2022. a knowledge-first approach to episodic memory. synthese 200(5). legate, julie anne. 2010. on how ‘how’ is used instead of ‘that’. natural language and linguistic theory 28(1). 121–134. 10.1007/s11049-010-9088-y. https://doi.org/10.1007/ s11049-010-9088-y. liddell, torrin & john kruschke. 2018. analyzing ordinal data with metric models: what could possibly go wrong? journal of experimental social psychology 328–348. liefke, kristina. 2023. two kinds of english non-interrogative, non-manner how-complements. in lukaszukasz jedrzejowski & carla umbach (eds.), non-interrogative subordinate wh-clauses oxford studies in theoretical linguistics, 24–62. oxford up. liefke, kristina & markus werning. 2024. diachronicity matters! how semantics supports discontinuism about remembering and imagining. topoi 1–23. 10.1007/s11245-024-10068-1. rosina, emil e. 2024. semantic experiments on ‘remember’ as gettier cases of memory. draft. https://lingbuzz.net/lingbuzz/008328/. rosina, emil e. & kristina liefke. 2024a. german ‘noch genau wissen’: uniform semantics, distinct effects. accepted for the proceedings of wccfl . rosina, emil e. & kristina liefke. 2024b. german ‘wie’ in memory reports. 10.17605/osf.io/v5peh. osf preregistration. rosina, emil e. & kristina liefke. 2024c. german ‘wie’ in memory reports intermediate evidence follow-up. 10.17605/osf.io/bm38w. osf preregistration. rosina, emil e. & kristina liefke. 2024d. german ‘wie’ in memory reports qud manipulation. 10.17605/osf.io/9bv3k. osf preregistration. sadock, jerrold m. & arnold m. zwicky. 1975. ambiguity tests and how to fail them. in j. kimball (ed.), syntax and semantics, vol. 4, 1–36. new york. stephenson, tamina. 2010. vivid attitudes: centered situations in the semantics of remember and imagine. semantics and linguistic theory (salt) 20. 147–160. 10.3765/salt.v0i20.2582. tulving, endel. 1972. episodic and semantic memory. in endel tulving & w. donaldson (eds.), organization of memory, 381–402. new york: academic press. umbach, carla, stefan hinterwimmer & helmar gust. 2022. german wie-complements: manners, methods and events in progress. natural language and linguistic theory 40. 307–343. 10.1007/s11049-021-09508-z. zhu, ziling & dorothy ahn. 2023. effects of instruction on semantic and pragmatic judgment tasks. in proceedings of elm 2, 322–330. ünal, ercenur & anna papafragou. 2018. relations between language and cognition: evidentiality and sources of knowledge. topics in cognitive science 12. 10.1111/tops.12355. proceedings of elm 3: 319-331, 2025 emil eva rosina and kristina liefke: experientiality markers in memory reports: a semantics-pragmatics puzzle. 331 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ investigating fragment usage with a gamified utterance selection task robin lemke* abstract. nonsentential utterances, or fragments, like a coffee, please! can often be used to communicate a propositional meaning otherwise encoded by a complete sentence (i’d like to order a coffee, please!). previous research focused mostly on the syntax and licensing of fragments, but the questions of why speakers use fragments and how listeners interpret them are still underexplored. i propose a simple gametheoretic account of fragment usage, which predicts (i) that listeners assign fragments the most likely interpretation in context and (ii) that speakers are aware of this and trade-off production cost and the risk of being misunderstood when choosing their utterance. using a corpus of production data, empirically founded and precise model predictions are generated. these predictions are evaluated with two experiments using a novel gamified utterance selection paradigm. the experiments suggest that, as predicted, speakers take into account both potential gains in efficiency and the risk of being misunderstood when choosing their utterance. keywords. ellipsis, fragments, game theory 1. introduction. instead of a complete sentence like (1a), speakers often use nonsentential utterances, or fragments (morgan 1973), for this purpose (1b). despite their reduced form, fragments can be meaning-equivalent to their fully sentential counterparts in an appropriate context. (1) [passenger asks conductor on the platform next to a waiting train:] a. does this train go to paris? b. to paris? while the syntax of fragments has been and is being extensively studied (see e.g. ginzburg & sag 2000, merchant 2004, reich 2007, weir 2014, ott & struckmeier 2016, lemke 2021), the question of why speakers actually choose to use a fragment is relatively underexplored (but cf. bergen & goodman 2015, lemke 2021). intuitively, fragments have the advantage of being shorter than sentences, which reduces the production effort for the speaker and the (syntactic) processing effort for the listener. this makes communication more efficient, as the same information is transmitted in less time, a tendency that also can be attributed to the gricean maxim of manner “be brief” (grice 1975). the downside of fragments is their vagueness: a single fragment can often be used to communicate meanings expressed by different complete sentences. for instance, the fragment in (1b) can communicate not only (1a), but also any of the questions in (2). due to this ambiguity, the listener might assign the fragment a different interpretation than intended by the speaker, or they might not be able to retrieve any interpretation at all. both of these outcomes would result in communication failure and require further clarifications, making communication less efficient. (2) a. are you traveling to paris? b. have you ever been to paris? *authors: robin lemke, saarland university (robin.lemke@uni-saarland.de). proceedings of elm 3: 447-459, 2025 c©2025 robin lemke published by the lsa with permission of the author(s) under a cc by license. 447 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ i pursue the hypothesis that speakers counterbalance the gain in efficiency provided by fragments with the risk of communication failure caused by their ambiguity when choosing how to encode the message. if the gain in efficiency outweighs the risk, the speaker will use the fragment; if it does not, a full sentence. this might explain the observation in lemke (2021) that fragments are preferred over sentences as answers to questions, but the opposite holds in discourse-initial contexts, where the uncertainty about their intended interpretation is probably greater. to formalize this idea, i use a simple game-theoretic model, which allows for including an explicit utterance cost structure and which predicts how likely it actually is to communicate a particular meaning in a context. such models, which have been widely applied to pragmatic phenomena, such as scalar implicature (franke 2009), reference (frank & goodman 2012, rohde et al. 2012) or the construction of social meaning (burnett 2017), predict speaker and listener behavior based on an explicit model of the utterance context. from the perspective of ellipsis research, this has the advantage of explaining the interpretation of elliptical utterances based on pragmatic inferences (i.e. relying on the same mechanisms as implicature generation), instead of having to assume a specific processing mechanism, as has been suggested by e.g. arregui et al. (2006) for vp ellipsis. in order to generate model predictions, i rely on a corpus of production data collected by lemke (2021). this approach differs from most previous empirical investigations of gametheoretic reasoning in pragmatics, which relied on controlled and balanced experimental setups in which objects are referred to by their shape or color with oneor two-word utterances (e.g. frank & goodman 2012, rohde et al. 2012, sikos et al. 2021). to my knowledge, my study is therefore the first attempt in game-theoretic pragmatics to derive model predictions from a relatively large, unbalanced, and diverse data set based on linguistic data produced by human subjects. this paper is organized as follows: section 2 presents the game-theoretic account of fragment usage and section 3 the data set on which model predictions are based. section 4 introduces the experimental design, before sections 5 and 6 present the results of the two experiments and section 7 summarizes the main results and conclusions. 2. a game-theoretic account of fragment usage. i formalize the idea that fragment usage results from a trade-off between the gain in efficiency and the risk of being misunderstood within a signaling game framework (franke 2009, frank & goodman 2012, sikos et al. 2021). the model predicts speaker and listener behavior based on a (possibly partial) mapping between a set of utterances and a set of messages, i.e. meanings the speaker could communicate. the components of the model are mostly equivalent to those in franke (2009), although i slightly adapt the terminology to better fit the inference from (possibly nonsentential) utterances to complete sentences. there is a set of messages m ∈ m comprising all of the meanings the speaker can communicate in a situation (e.g. (1a) and (2) in the train example), and a set of utterances u ∈ u that they can use for this purpose. the speaker’s task is to send the utterance which is optimal to communicate the intended message. the listener receives this utterance and performs an interpretation action a ∈ a, i.e. they assign the utterance a meaning. since m contains all possible meanings, a = m . the set of utterances u contains (i) sentences, which are equivalent to the m ∈ m 1 and 1this is a simplification, because there are also meaning-equivalent sentences. i address this concern by pooling the production data to sentences differing in meaning as discussed in section 3, but it would also be possible to include all of the synonym sentences in u . proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 448 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (ii) all fragments which can be derived from the m ∈ m by grammatically licensed omission.2 i consider only fragments that are constituents or sequences thereof (e.g. to paris?, this (one) to paris?, but not, for example, the heads of a dp and a pp only (e.g. this to?). the reason for this is that fragments are licensed by focus or givenness (e.g. merchant 2004, reich 2007) and these concepts apply generally to constituents. having established m and u , it is necessary to determine whether an utterance u can be used to communicate a message m. in his account of implicature, franke (2009) relies on truth conditions, but fragments cannot be considered true or false before having been enriched to a complete proposition. given the computation of u described above, i rely on whether u can be derived from m by grammatically licensed omission. following the notation in franke (2009), assume a denotation function [[·]], which returns 1 if u can be derived from m and 0 otherwise. finally, not all messages are necessarily equally likely in a situation. this might be due to world knowledge (people are more likely to ask the conductor whether the train goes to paris than whether they have been to paris) or statistical knowledge about language use. this potential difference in likelihood is captured by a prior probability distribution pr(m). based on these components, following franke (2009), chains of mutual reasoning about the behavior, preferences and beliefs of the interlocutors are initialized: one starting from a “literal” speaker s0 and one starting from a “literal” listener l0. since i model the speaker’s perspective in my experiments, i focus on the basic listener model, which the speaker should consider when selecting their utterance. given a message u sent by the speaker, l0 calculates the likelihood of each m ∈ m given u. as equation 1 shows, this posterior p(m|u) is determined as the ratio of pr(m) and the probability mass of other possible interpretations m′ of u. this reweighs the prior probability among those messages from which u could have been derived. based on the resulting probability distribution, the listener selects the most likely message as the interpretation of u. as i discuss in section 4.2, the reason for this is that successful communication is rewarded with a payoff and maximizing l0p(m,u) also maximizes the payoff. l0(m,u) = pr(m)× [[u]]m∑ m′ pr(m′)× [[u]]m′ (1) a speaker who takes listener behavior into account (a s1 speaker in franke’s terminology) will constrain their utterance choice based on the likelihood that the listener will interpret it as intended. since u has been derived by omission from m , sentential utterances fully disambiguate between messages. therefore, from the perspective of successful communication they should be always at least as ideal as a fragment. however, sentences come with an additional production cost, which might result in them being less ideal for the speaker.3 since it is a priori unclear how to quantify utterance cost and its relationship to the gain in efficiency, i use an explicit cost structure in my experiments to ensure that sending fragments is cheaper than sending sentences. 2note that i follow most of the theoretical literature cited above and treat fragments as grammatically well-formed, but there are diverging views: bergen & goodman (2015) assume that fragments are ungrammatical but can still be suitable means of communication if the listener can retrieve their meaning. the account i propose is in principle independent from the question of whether fragments are grammatical, but assuming ungrammaticality would require to include ungrammatical fragments in u , which i do not. 3of course, unnecessary redundancy might be dispreferred by the listener (schäfer et al. 2021), but since they have no control over which utterance the speaker sends, this does not need to be modeled. proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 449 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 3. data set. testing the predictions of the account described in section 2 empirically requires realistic estimates of the model’s components: (i) the set of messages m , (ii) the set of utterances u , (iii) the mapping between m and u , and (iv) the prior probability distribution over messages pr(m). i calculate these parameters from a data set originally collected by lemke (2021).4 lemke (2021) used a crowd-sourced written production task to elicit utterances with script-based context stories like (3). the participants were instructed to produce the most natural utterance in that situation by the person specified in the prompt annika to jenny: (3) annika and jenny want to cook pasta. annika has put a pot with water on the stove. then she has turned the stove on. after a few minutes, the water has started to boil. annika to jenny: the stories were based on event chains extracted from the descript corpus of script knowledge (wanzare et al. 2016). for each story, which represents a different script-based scenario, about 100 responses were collected. the answers (e.g. (4a)) were preprocessed into abstract representations like (4b), where each word corresponded to a constituent that could be freely omitted in german. this ensures that all fragments derived from these representations are grammatical, i.e. constituents or sequences thereof, as discussed in section 2.5 (4) a. pour the pasta into the pot! b. pour pasta pot.goal the set of unique representations for each scenario (e.g. cooking pasta) was taken to represent m in this scenario. since german has free word order, all of the words in the representations like (4b) were ordered alphabetically. this ensures that the synonym pour pasta pot.goal and pour pot.goal pasta are not treated as different meaning representations. the likelihood of each representation within the 100 preprocessed responses determines pr(m) in this scenario. u is the union of m and all fragments that can be derived by grammatical omission from all m ∈ m : for a message m, this includes the corresponding sentence, all of its constituents (i.e. the individual “words” in the abstract representation) and all possible combinations of thereof. (5) exemplifies this at the example of (4b). (5) pour pasta pot.goal, pasta pot.goal, pour pot.goal, pour pasta, pour, pasta, pot.goal finally, [[u]]m determined for each u ∈ u and each m ∈ m whether u could be derived from m by omission. applying equation 1, the likelihood of communicative success l0(m|u) was then calculated for each m ∈ m and u ∈ u . 4. experimental approach. the experiments presented below were designed to test the hypothesis that the usage of fragment is conditioned by a trade-off between the gain in efficiency and the risk of being misunderstood. therefore, fragments are expected to be more strongly preferred if (i) their cost relative to a sentence is lower, and if (ii) the likelihood of communicative success is relatively high. 4the data set was collected in german, but i present the context stories in english for convenience. 5for details on the preprocessing procedure, see lemke (2021; 202–206). proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 450 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ i tested this with a pseudo-interactive gamified utterance selection task inspired by rohde et al. (2012). in my experiments, the participant took the speaker role and chose between different sentential and fragment utterances to communicate a message determined by the experiment. the listener was simulated to behave according to model predictions: for sentences, they always selected the correct interpretation and for fragments, they maximized l0p(m|u). an explicit cost structure ensured that fragments were cheaper than sentences. 4.1. materials: context story, messages, and utterances. the stimuli were derived from the data set by lemke (2021) introduced above. in each trial, the participant saw a context story, three messages and six utterances, as shown in figure 1. the context stories were based on those used by lemke (2021), but the roles of the characters were assigned to the participant and their (simulated) partner. the messages were presented as states of affairs that the participant might want to communicate to their partner. the message the participant had to communicate was determined by the experiment depending on the experimental condition (see section 4.2) and highlighted in blue. among the six utterances, three were sentences, each unambiguously encoding one of the messages. one of the three fragments (into the water in figure 1) is ambiguous and can communicate two of the messages. these two messages were selected from the production data so that one of them has a higher l0p(m|u) given the ambiguous fragment than the other one. if participants base their utterance choice on the likelihood of communicative success, they should be more likely to use a fragment when asked to refer to the more likely message than when referring to the less likely one, because a listener maximizing l0p(m|u) will choose the more likely interpretation more often. the second fragment (in the example: on the table) always referred unambiguously to the third message. this establishes a baseline for the rate of fragment usage when there is no risk of misunderstanding in the experimental setup. the third fragment (the recipe!) had no corresponding message and served as a control to exclude inattentive participants who randomly click on any of the utterances. all messages and utterances were based on the abstract representations contained in the production data for the respective scenario. 4.2. conditions, utterance cost and expectations. there were three conditions differing in the message to be communicated. in the critical condition, the most likely message given the ambiguous fragment was highlighted. in the competitor condition, the less likely message compatible with the ambiguous fragment was highlighted. in the unambiguous condition, the message to which the unambiguous fragment refers was to be communicated. note that this fragment was unambiguous with respect to the three messages displayed in the experimental setup, though not necessarily in the production data set.6 as table 1 shows, on average, the ambiguous fragment had a higher l0p(m|u) in the critical condition than in the competitor condition within each scenario (see table 1). as the ranges show, there was some overlap in the probabilities, but within each scenario, the message in the critical condition was always more likely. production cost was implemented with a system of virtual coins. in experiment 1, participants received a starting balance of 500 coins and the cost for sending a fragment (30 coins) was lower than that for sentences (100 coins). successful communication was rewarded with 120 6in some stimuli, it was also unambiguous in the production data, as the maximal l0p(m|u) of 1 in table 1 shows, but this was not always the case. proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 451 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: screenshot of experiment 1 after revealing the messages and utterances. the experiment was conducted in german, but has been translated here for convenience. condition lowest p(m|u) highest p(m|u) mean p(m|u) critical 0.12 0.69 0.38 competitor 0.03 0.22 0.09 unambiguous 0.15 1.0 0.75 table 1: range of l0(m|u) probabilities and means by conditions coins.7 these quantities were set based on the expected utility (eu) (franke 2009) of the utterances in each condition. in the experiment, eu can be calculated as shown in equation 2 by subtracting the cost c from the product of the likelihood of successful communication l0p(m|u) and the utility u(m,u), i.e. the payoff in case of success. eu(u,m) = l0p(m|u)× u(m,u)− c (2) 7see section 6 for the cost structure used in experiment 2. proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 452 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ since the eu of an (unambiguous) sentence is eu(u,m) = 1 × 120 − 100 = 20 and fragments yield a higher eu if their l0p(m|u) > 0.42, the model predicts an absolute preference for fragments beyond this threshold. given the data in table 1, in the range of l0p(m|u) it should be possible to observe differences in the ratio of fragments produced. on average, the fragment ratio is expected to be highest in the unambiguous condition and lowest in the competitor condition, with the critical condition in between. 4.3. procedure. each of the two experiments was completed by 60 self-reported native speakers of german recruited on the crowd-sourcing platform prolific. the participants were rewarded with £2.67 (the equivalent of the german minimum wage of 12.41c for the projected duration of 15 minutes). the materials were distributed across three lists, so that each participant saw each of the 15 stimuli in one of three conditions, with each condition appearing equally often. stimuli were presented in individual fully randomized order. 8 after giving informed consent by marking a checkbox, reading the instructions and providing their prolific id, participants were asked to enter a nickname, which should not be their real name and which would not be published, to make the experiment appear more interactive. they were then asked to wait for the connection to their partner to be established. the waiting time was randomly sampled from an array of durations between 10 and 25 seconds. after this, the participants were notified that a partner had been found, informed about their partner’s name (randomly sampled from an array of names), and asked to begin a practice phase of two (experiment 1) or three (experiment 2) trials.9 the setup in the practice phase was identical to the main experiment in terms of the number of messages and utterances, as well as the costs. the only difference was that each of the three fragments referred unambiguously to one of the messages. after the practice phase, the score was reset to the initial amount and participants were informed in advance about this. each trial began with a display of the context story only. participants were told to press the space bar to display each of the other elements: the messages, the highlighted message, and the utterances. this intended to make subjects read all of the messages before seeing which one would be highlighted, which is critical to realize that the ambiguous fragment is ambiguous. after selecting an utterance, participants received feedback on the partner’s interpretation. as anticipated above, the partner always maximized l0p(m|u) when selecting an utterance: for correct sentential utterances and fragments in the critical condition (where the message had a high l0p(m|u)) this results always in correct interpretations. fragment utterances in the (low l0p(m|u)) competitor condition were never successful, as maximizing l0p(m|u) led to selecting the more likely interpretation. if communication was successful, a notice “⟨partner⟩ understood you correctly” appeared in the center of the screen, otherwise it read “⟨partner⟩ understood something else”. additionally, the message selected by the fake partner was highlighted in green (success) or red (failure) to indicate what the partner understood. if the participant had selected the attention check fragment, which 8the experiment was followed by a questionnaire to evaluate the experiment, which included questions about the participants’ affinity to games and self-assessed likelihood of taken risks to investigate potential individual differences. since there were no theoretically interesting correlations between the answers to these questions and behavioral data, the questionnaire is not discussed further here. 9since there were almost no errors in the practice phase, it was shortened to two trials in experiment 2. proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 453 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 0.00 0.25 0.50 0.75 1.00 competitor critical unambiguous f ra g m e nt r a ti o competitor critical unambiguous 0.0 0.2 0.4 0.6 0.8 0.00 0.25 0.50 0.75 1.00 l0 p(m|u) f ra g m e nt r a ti o competitor critical unambiguous figure 2: the left facet shows the ratio of fragments across the three conditions, the right facet illustrates the continuous effect of l0(m|u) on the ratio of fragments with a loess smooth (span=1). did not refer to any of the three messages, the text field “⟨partner⟩ is not sure” was highlighted. to simulate the partner’s reasoning, a delay was introduced between the participant’s sending the message and the feedback. the delay time was randomly sampled from an array of values between 0.8 and 4.4 seconds. since the participant was told that their partner could see the complete screen throughout the trial, and consequently have read all of the messages and utterances in the meantime, quick responses to the participants’ choice seemed natural. 5. experiment 1. experiment 1 was conducted using the cost structure described in section 4.2. due to an error in preparing the stimulus table, which resulted in the target fragment being unambiguous, this stimulus had to be discarded before the analysis, leaving 14 stimuli for analysis. one participant, who provided more utterances incompatble with the target message (n=8) than compatible ones (n=6), was also excluded from further analysis. 5.1. results. the participants were very accurate: only in three trials, the distractor was chosen and in 16 trials an utterance that could not communicate the highlighted message. these 19 trials (2.3% of the data) were excluded before the analyses reported in what follows. the responses across the three conditions and as a function of l0p(m|u) are summarized in figure 2. overall, the participants had a strong preference for complete sentences: in the ambiguous conditions, the fragment was selected in only about 12-13% of the trials, and even in the unambiguous condition, participants preferred the fragment only in half of the trials. since the game-theoretic account predicts a gradual effect of an utterance’s expected utility on speakers’ choices, and there are notable differences in l0p(m|u) between the scenarios tested (see table 1), in the statistical analyses, the probability of successful communication was used as a continuous predictor in the statistical analyses rather than investigating categorical contrasts between the proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 454 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ three conditions. descriptively, the right facet of figure 2 supports the prediction that a higher likelihood of successful communication with fragments increases the fragment ratio. the data were analyzed with mixed effects logistic regressions (bates et al. 2015) in r (r core team 2024; version 4.4.1) using a backward model selection procedure. the models predicted a binary dependent variable ellipsis (sentence/fragment) from the probability (l0p(m|u)) of the fragment and the scaled and centered position of the trial in the experiment. the full model contained the maximal random effects structure supported by the data, which consisted of bysubject and by-item random intercepts and by-item random slopes for probability.10 starting from this full model, fixed effects that did not significantly improve model fit were subsequently removed. this was determined using likelihood ratio tests conducted with the anova function (r core team 2024). the final model contained only a significant main effect of probability (χ2 = 6.5, p < 0.05), which suggests that, as predicted by the game-theoretic account, the ratio of fragments increases with their l0p(m|u). this finding is consistent with the game-theoretic approach. however, the left facet of figure 2 suggests that there may be no difference between the ambiguous (critical and competitor) conditions, despite the higher l0p(m|u) in the critical condition predicting a higher fragment ratio. consequently, the probability effect might be driven by the unambiguous condition alone. this aligns with the account’s predictions, because the fragments in the unambiguous condition had a higher average l0p(m|u). however, this does not provide genuine evidence for probabilistic rational reasoning, as participants may have simply avoided ambiguity in the experiment. therefore, i conducted a further regression analysis using only the data for the ambiguous conditions. if participants relied on game-theoretic reasoning, the effect of probability should be replicated in this subset. note that the opposite result would not prove that participants did not rely on game-theoretic reasoning, as they might it and be but less likely to assume risks than predicted by the current model. the analysis followed the same procedure as the main analysis reported above in this section. however, there was no significant main effect of probability (χ2 = 0.01, p > 0.9), failing to confirm the game-theoretic prediction of a gradual effect. 5.2. discussion. the results of experiment 1 are in line with the prediction of the gametheoretic account that speakers choose ambiguous fragments more often if the risk of misunderstanding is reduced due to a high l0p(m|u). this gradual effect was not confirmed when excluding the unambiguous condition, which could have at least two reasons. first, fragments had the highest l0p(m|u) in this condition, so that the preference for using fragments might simply be stronger. second, if participants did not rely on game-theoretic reasoning and prior message probabilities, they could simply avoid any expression which is ambiguous in the context of the experiment. while the former explanation would support the game-theoretic account, the second would provide an explanation for the results which does not require probabilistic pragmatic reasoning. another result of the experiment is the overall relatively low fragment ratio. even for some fragments with a l0p(m|u) of 1 in the unambiguous condition, participants selected the fragments only in 30–50% of the trials. this might be at least partially due to the cost structure, which guaranteed participants a net benefit of 20 coins per trial when using only unambiguous sentences. 10ellipsis ∼ probability * position + (1 | subject) + (1 + probability | item) proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 455 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 0.00 0.25 0.50 0.75 1.00 competitor critical unambiguous condition f ra g m e nt r a ti o competitor critical unambiguous 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 l0 p(m|u) f ra g m e nt r a ti o competitor critical unambiguous figure 3: the left facet shows the ratio of fragments across the three conditions, the right facet illustrates the continuous effect of l0(m|u) on the ratio of fragments with a loess smooth (span=1). when looking at the individual participants’ data, it turns out that 15 of them correctly completed all trials by selecting only sentences. this suggests that the cost structure used in experiment 1 biased participants toward adopting a risk-avoiding strategy, which might have obscured gradual differences between the ambiguous conditions. 6. experiment 2. experiment 2 addressed the concern that the low fragment ratio in experiment 1 might have occurred due to a cost structure favoring a risk-avoiding strategy which made about 25% of the participants to select sentences only. to address this, the cost structure was modified to reduce the benefits obtained in experiment 1 through this strategy. 6.1. materials and cost structure. the materials used in experiment 2 were identical to experiment 1, except that the error resulting in the exclusion of one experimental item was resolved. to make a sentence-only strategy less rewarding, the utterance costs and the starting balance were adjusted: the reward for successful communication was reduced to 100 coins (instead of 120), the cost of sentences increased to 130 (instead of 100) and the cost of fragments increased to 40 (instead of 30). unlike in experiment 1, selecting sentences now resulted in a loss of 30 coins per trial instead of a benefit of 20 and the lower starting balance should increase the motivation to keep the score high (even though a negative score had no consequences for the subjects’ reward and they were informed about this in advance). 6.2. procedure. the procedure was identical to experiment 1. 60 further participants, who did not participate in experiment 1, were recruited on prolific and rewarded with £2.67. 6.3. results. figures 3 shows the ratio of fragments by condition and as a function of l0p(m|u). the overall fragment ratio was higher than in experiment 2 (35.01% instead of 25.32%), which proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 456 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ predictor est. std. error χ2 p-value intercept 2.14 0.36 17.71 < 0.001 probability −2.56 0.67 9.23 < 0.01 experiment −0.50 0.12 18.02 < 0.001 table 2: fixed effects in the final model for the joint analysis of experiments 1 and 2. was also reflected in participants’ individual behavior: unlike in experiment 1, only three participants followed a sentence-only strategy, which resulted in a score of -150. this suggests that the smaller but secure reward that this strategy returned in experiment 1 was (at least partially) responsible for the higher relatively low fragment usage in experiment 1). besides the higher fragment ratio, the pattern looks similar to experiment 1: fragment ratio seems to increase as a function of l0p(m|u), but there is no obvious difference between the ambiguous conditions. i first conducted an analysis in parallel with experiment 1 using mixed effects logistic regressions (bates et al. 2015; version 1.1-35.3) in r (r core team 2024; version 4.4.1). the full model contained by-subject and by-item random intercepts and by-item slopes for probability, as well as fixed effects for probability, position (scaled and centered) and their interaction. the overall pattern was identical to experiment 1: in the analysis of the complete data set, there was a significant main effect of probability (χ2 = 11.04, p < 0.001). as in experiment 1, this effect was not significant within the ambiguous conditions only (χ2 = 0.4, p > 0.5). to quantify the difference between the experiments, i then conducted a joint analysis of the data from both studies. the procedure was identical to the regression analyses reported above, except that i included an additional predictor experiment (sum-coded, exp.1 = −0.5, exp.2 = 0.5) and its interactions with probability (numeric) and position (numeric, scaled and centered). a significant main effect of experiment would statistically confirm the higher fragment ratio in experiment 2 and an experiment:probability interaction would show whether the effect of probability is also stronger. the full model contained main effects for all three predictors and all interactions between them, as well as by-items random intercepts and random slopes for probability. the final model (see table 2) contains a main effect of probability, which had been also found in the analysis of the individual experiments (χ2 = 9.23, p < 0.01), a main effect of experiment (χ2 = 18.02, p < 0.001). the latter effect confirms the intuition that the fragment ratio is higher in experiment 2. the interaction between both predictors is not significant (χ2 = 2.14, p > 0.1). 6.4. discussion. experiment 2 investigated whether the low fragment ratio in experiment 1 was (partially) due to an overall bias toward sentences in order to avoid risks. this could have obscured gradual differences between the ambiguous conditions, which would have provided genuine evidence for pragmatic reasoning in fragment usage. the significant difference in fragment ratio between both experiments, confirmed by the joint analysis, suggests that the cost structure had an effect: if fragments yield a higher potential benefit, their ratio is increased, even if the likelihood of communicative success remains constant. however, it is also clear that the stronger preference for sentences in experiment 1 was not the only cause for the lack of a gradual effect within the ambiguous conditions, since this effect was also not found in experiment 2. proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 457 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 7. general discussion. i conducted two pseudo-interactive utterance selection experiments to investigate whether the usage of fragments follows the predictions of a game-theoretic account. unlike previous studies in the field, i did not use tightly restricted and balanced sets of utterances and messages; instead, i constructed the stimuli based on a diverse, crowd-sourced data set, with a much larger variety of messages and utterances, as well as unbalanced priors over messages. the results are overall in line with the trade-off between reducing production cost and maximizing the likelihood of successful communication predicted by the game-theoretic account: in the main analyses of both experiments, participants preferred fragments more often when the model predicted a high likelihood of getting the message across. they also selected more fragments in experiment 2, where the modified cost structure increased the gain in efficiency as compared to experiment 1. however, the effect observed in the complete data set was not found when examining only the data from the ambiguous conditions. experiment 2 investigated whether it had been obscured by the bias toward using sentences in experiment 1, but this was not the case. therefore, at this point it cannot be ruled out that participants simply avoided ambiguity in the experimental setup, which does not require pragmatic reasoning. a possible reason for the absence of the gradual effect might be the method used to collect the production data set. since participants provided only a single most likely utterance, less likely utterances are probably underrepresented in the data set. this could result in low l0p(m|u) messages being relatively predictable, which may have biased subjects to prefer sentences more often. taken together, the experiments show how a game-theoretic account of language production and interpretation can also be applied to more realistic and diverse communication situations than the previous tightly controlled studies on reference. they also suggest that the choice between an elliptical and a sentential utterance depends on a trade-off between production cost and the likelihood of communicative success, which can be explicitly captured by a game-theoretic model of pragmatic inference. future research could explore the extent to which this reasoning is required only for interpreting discourse-initial fragments tested in the experiments presented here, or whether it also extends to other types of ellipsis. references arregui, ana, charles clifton, lyn frazier & keir moulton. 2006. processing elided verb phrases with flawed antecedents: the recycling hypothesis. journal of memory and language 55(2). 232–246. 10.1016/j.jml.2006.02.005. bates, douglas, martin mächler, ben bolker & steve walker. 2015. fitting linear mixed-effects models using lme4. journal of statistical software 67(1). 1–48. 10.18637/jss.v067.i01. bergen, leon & noah goodman. 2015. the strategic use of noise in pragmatic reasoning. topics in cognitive science 7(2). 336–350. 10.1111/tops.12144. burnett, heather. 2017. sociolinguistic interaction and identity construction: the view from gametheoretic pragmatics. journal of sociolinguistics 21(2). 238–271. 10.1111/josl.12229. frank, michael. c. & noah. d. goodman. 2012. predicting pragmatic reasoning in language games. science 336(6084). 998–998. 10.1126/science.1218633. franke, michael. 2009. signal to act: game theory in pragmatics: universiteit van amsterdam dissertation. proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 458 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ ginzburg, jonathan & ivan a. sag. 2000. interrogative investigations: the form, meaning, and use of english interrogatives. stanford, california: csli publications. grice, h. paul. 1975. logic and conversation. in p. cole & j.l. morgan (eds.), syntax and semantics, vol. 3, 41–58. new york: academic press. lemke, robin. 2021. experimental investigations on the syntax and usage of fragments (open germanic linguistics 1). berlin: language science press. merchant, jason. 2004. fragments and ellipsis. linguistics and philosophy 27(6). 661–738. 10.1007/s10988-005-7378-3. morgan, jerry. 1973. sentence fragments and the notion ‘sentence’. in braj b. kachru, robert lees, yakov malkiel, angelina pietrangeli & sol saporta (eds.), issues in linguistics. papers in honor of henry and renée kahane, 719–751. urbana: university of illionois press. ott, dennis & volker struckmeier. 2016. deletion in clausal ellipsis: remnants in the middle field. upenn working papers in linguistics 22(1). 225–234. r core team. 2024. r: a language and environment for statistical computing. vienna, austria. reich, ingo. 2007. toward a uniform analysis of short answers and gapping. in kerstin schwabe & susanne winkler (eds.), on information structure, meaning and form, 467–484. amsterdam: john benjamins. 10.1075/la.100.25rei. rohde, hannah, scott seyfarth, brady clark, gerhard jaeger & stefan kaufmann. 2012. communicating with cost-based implicature: a game-theoretic approach to ambiguity. in proceedings of the 16th workshop on the semantics and pragmatics of dialogue, 107–116. schäfer, lisa, robin lemke, heiner drenhaus & ingo reich. 2021. the role of uid for the usage of verb phrase ellipsis: psycholinguistic evidence from length and context effects. frontiers in psychology 12. 10.3389/fpsyg.2021.661087. sikos, les, noortje j. venhuizen, heiner drenhaus & matthew w. crocker. 2021. reevaluating pragmatic reasoning in language games. plos one 16(3). e0248388. 10.1371/journal.pone.0248388. wanzare, lilian d. a., alessandra zarcone, stefan thater & manfred pinkal. 2016. descript: a crowdsourced corpus for the acquisition of high-quality script knowledge. in proceedings of lrec 2016, 3494–3501. portoroz, slovenia. weir, andrew. 2014. fragments and clausal ellipsis: university of massachusetts amherst dissertation. proceedings of elm 3: 447-459, 2025 robin lemke: investigating fragment usage with a gamified utterance selection task. 459 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ modeling the prompt in inference judgment tasks julian grove & aaron steven white* abstract. we show that when analyzing data from inference judgment tasks, it can be important to incorporate into one’s data analysis regime an explicit representation of the semantics of the natural language prompt used to guide participants on the task. to demonstrate this, we conduct two experiments within an existing experimental paradigm focused on measuring factive inferences, while manipulating the prompt participants receive in small but semantically potent ways. in statistical model comparisons couched within the framework of probabilistic dynamic semantics, we find that probabilistic models structured, in part, by the semantics of the prompt fit better to data collected using that prompt than models that ignore the semantics of the prompt. keywords. presupposition, factivity, dynamic semantics, probabilistic models 1. introduction. when collecting inference judgments in formal experiments, it is common for trials to consist of the following pieces: (i) a target linguistic expression whose inferential affordances one is interested in measuring; (ii) a description of some context which the expression might be used in; (iii) a natural language prompt guiding participants on the relevant task; and (iv) a response instrument, such as an ordinal or slider scale. when analyzing the data from such experiments, one generally incorporates some representation of components (i), (ii), and (iv); but it is relatively rare to incorporate a representation of component (iii)—likely because this component is typically constant across all experimental items.1 we show that it can be important to incorporate an explicit representation of the semantics of natural language prompts used in inference judgment experiments into one’s data analysis regime. to demonstrate this, we fix components (i), (ii), and (iv) of an experimental paradigm focused on measuring factive inferences—such as the inference from (1a) to (1b)—and manipulate component (iii)—the natural language prompt—in small but semantically potent ways. (1) a. jo {loves, doesn’t love} that mo left. b. mo left. we collect data using two such manipulated prompts and show, through statistical model comparison, that probabilistic models structured, in part, by the semantics of the prompt fit better to data collected using that prompt than models that ignore its semantics. in section 2, we describe how we incorporate natural language prompts into statistical models using the framework of probabilistic dynamic semantics (pds; grove & white 2024a,b). we then review, in section 3, the experimental paradigm and prior models within pds of data collected *we thank the organizers of elm 3, as well as four anonymous reviewers. authors: julian grove, university of rochester (julian.grove@gmail.com) & aaron steven white, university of rochester (aaron.white@rochester.edu). 1one might conceive of the way that hypotheses are modeled in the natural language inference (nli) literature (cooper et al. 1996, dagan, glickman & magnini 2006, maccartney 2009, et seq.) as an exception; but most nli models assume a single ground truth label aggregated from possibly multiple participants responses (cf. gantt, kane & white 2020), rather than modeling the distribution of reponses themselves, as we do here. proceedings of elm 3: 176-187, 2025 c©2025 julian grove and aaron steven white published by the lsa with permission of the author(s) under a cc by license. 176 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ under that paradigm. in section 4, we describe our two modifications to this paradigm and the two corresponding experiments. in section 5, we describe our model implementations and report model comparisons before concluding in section 6. 2. probabilistic dynamic semantics. we couch our modeling within the recently developed probabilistic dynamic semantics (pds) framework (grove & white 2024b). one of the core ideas we implement within this framework is that an inference judgment task may be characterized as a kind of discourse: a target sentence and its context of interpretation may be analyzed as updates to a common ground; meanwhile, the prompt may be regarded as contributing a new question under discussion (qud; ginzburg 1996, roberts 2012). the upshot is that once we fix these formal components, we need only make linking assumptions—i.e., assumptions about how the probability distribution over possible answers to the qud (given a common ground, appropriately updated) should manifest as a particular distribution of responses, given some response instrument. our approach is therefore analogous to standard analysis regimes, insofar as one is always making linking assumptions via one’s choice of statistical model (see jasbi, waldon & degen 2019), but it has the additional benefit of providing an explicit interface between the formal semantics of an experimental item and such linking assumptions. we provide certain crucial details below. for a full introduction to pds, see grove & white 2024b. 2.1. distributions on discourse constructs. in pds, an ongoing discourse is represented by a map from an input state to a probability distribution over output states. states are simply tuples of metalinguistic parameters; these include, e.g., the common ground (stalnaker 1978) and the qud, as well as other conversationally relevant features of discourse (e.g., possible antecedents for elliptical constructions). one can view the state as akin to the context state of farkas & bruce (2010); our state, however, is less constrained in principle, in the sense that it is compatible with arbitrary representations of the context.2 notating the types of these state tuples ‘σ’, ‘σ′’, etc., an ongoing discourse is of type σ → pσ′, where pσ′ is the type of probability distributions over states of type σ′. (more generally, pα is the type of probability distributions over values of type α.) in words, a discourse is a map from input states to probability distributions over output states. to perform updates of various types to some ongoing discourse, we use two operators— bind and return—which allow probability distributions to be sampled (bind) and ordinary, nonprobabilistic values to be returned as degenerate distributions over those values (return).3 to illustrate, suppose we have some categorical distribution mammal : pe on mammals (where e is the type of entities). we can represent a distribution on mammals’ mothers as: x ∼ mammal mother(x) 2furthermore, as grove & bernardy (2023) show, typed frameworks for probabilistic semantics can be used to encode rational speechs acts (rsa) models, as laid out in frank & goodman 2012, goodman & frank 2016. frameworks like pds diverge from common implementation assumptions in such settings, however, in the sense that they have strong commitments to computational purity that often do not hold in practical applications of rsa (for discussion, see grove & bernardy 2023, as well as grove & white 2024b). 3importantly, these operators conform to the monad laws. see grove & bernardy 2023, grove & white 2024b for why this fact is useful. proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 177 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ here, mother : e → e maps an entity to its mother. meanwhile, x is sampled from mammal using bind, and mother(x) is then returned, as notated by the orange box. the idea is that this distribution assigns to any given entity y the probability ∑ x∈mother−1(y) pmammal(x) (where mother−1 : e → e → t is the inverse of the mother function and pmammal(x) is the probability that mammal assigns to x). in a similar fashion, given two discourses d1 : σ → pσ′ and d2 : σ′ → pσ′′, we can invoke bind to sequence d1 with d2 and obtain a new discourse d3 : σ → pσ′′: d3 = λs. ( s′ ∼ d1(s) d2(s ′) ) somewhat metaphorically: given a starting state s, we sample an output state from d1(s) and apply d2 to this output state. moreover, generalizing the previous example, given an input state s, the probability pd3(s)(s ′′) that d3(s) assigns to an output state s′′ is computed as: pd3(s)(s ′′) = ∑ s′ pd2(s′)(s ′′) ∗ pd1(s)(s ′) in words, sequencing two maps from input states to probability distributions over output states involves doing pointwise multiplication of the probabilities assigned to intermediate states by the first map and those assigned to an output state by the second map and then summing the result. finally, we represent the common ground as a probability distribution over indices representing information about possible worlds, along with certain linguistic parameters (which, for current purposes, we keep fairly open ended); we refer to the type of indices as ‘ι’, so that common grounds themselves are distributions of type pι. given some index i : ι, we refer to the “world” of i as ‘wi’ and to what we call the “context” of i as ‘ci’; contexts are what we use to encode useful linguistic information. the idea is that worlds w determine, e.g., how tall a particular individual is, while contexts c determine, e.g., the vague threshold of height past which one’s height is considered tall. thus while a world-sensitive function such as height(wi) : e → r intuitively says something about how tall an entity is at an index i, the adjective’s threshold dtall(ci) : r only provides information about an aspect of the lexical semantics of tall at that index, viz., what is required of an entity to count as tall. 2.2. making an assertion. similar to discourses, the meanings of expressions are both probabilistic and dynamic. accordingly, we represent them as functions of type σ → p(α × σ′): given an input state, the meaning of an expression produces a probability distribution over pairs of ordinary meanings of type α and possible output states. furthermore, given a sentence whose probabilistic dynamic meaning ϕ is of type σ → p((ι → t)× σ′), we can represent an assertion of that sentence as a discourse which updates the common ground. specifically, we have a function assert with the following type: assert : (σ → p((ι → t)× σ′)) → σ → pσ′ given such a ϕ, assert(ϕ) is a discourse of type σ → pσ′ representing an assertion of ϕ. for any input state s : σ, assert(ϕ)(s) samples a proposition p together with an output state s′ from ϕ and simply updates the common ground of s′ with p; in particular, it gives back a new state s′′ just proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 178 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ like s′, except that the common ground of s′′ (cg(s′′)) updates the common ground of s′ (cg(s′)) with p. ultimately, assertions modify an ongoing discourse so that its probability distribution over output states involves common grounds in which the proposition returned by ϕ has been observed to hold true (see grove & white 2024b for details). 2.3. asking a question. we follow a categorial tradition by analyzing questions as denoting— given an index—sets of true short answer meanings (hausser & zaefferer 1978, hausser 1983, xiang 2021; cf. karttunen 1977, groenendijk & stokhof 1984). for simplicity (but not by necessity), we assume that all questions are degree questions: given index, they denote sets of degrees of type r (real numbers). questions therefore have probabilistic dynamic meanings of type σ → p((r → ι → t)× σ′). asking a question is a matter of pushing a question meaning onto the qud stack (roberts 2012, farkas & bruce 2010). reflecting this, there is a function ask with the following type: ask : (σ → p((r → ι → t)× (σ′ 1 × δ × σ′ 2))) → σ → p(σ′ 1 × ((r → ι → t)× δ)× σ′ 2) given a probabilistic dynamic question meaning κ and an input state s, ask(κ)(s) samples a pair of a question meaning q : r → ι → t and a state s′ : σ′ 1 × δ × σ′ 2 from κ(s) and then updates the qud stack of s′ (something of type δ) by pushing q onto it. thus it returns a new output state of type σ′ 1 × ((r → ι → t)× δ)× σ′ 2. 2.4. responding to a question. pds also models responses to questions; at any point in an ongoing discourse, one can respond to the qud at the top of the current qud stack based on one’s prior knowledge. concretely, a given responder has some background knowledge bg : pσ constituting a prior distribution over starting states. the responder uses this prior, in conjuction with the interim updates to the discourse, to derive a probability distribution over answers to the qud. this answer distribution is gotten by retrieving the qud of any given state s′—resulting in a distribution over quds—and then taking the maximum degree of which it is true at an index sampled from the common ground—resulting, finally, in a distribution over degrees.4 2.5. linking assumptions. in practice (e.g., in the setting of a formal experiment), an answer needs to be given using a particular testing instrument. we assume that a given testing instrument may be modeled by a family f of distributions representing the likelihood, which is then fixed by a collection φ of nuisance parameters. thus we may define a family of response functions, parametric in the testing instrument (i.e., likelihood function), each of which takes a distribution bg representing one’s background knowledge, along with an ongoing discourse m: respondfφ:r→pρ : pσ → (σ → pσ′) → pρ for a fixed likelihood function fφ mapping any given real number answer onto a distribution over possible responses of type ρ (for some ρ), the response function takes a distribution representing background knowledge and a discourse to produce a response distribution. it does this by composing the discourse with background knowledge, as above, and then obtaining the maximum degree answer to the current qud, before applying the likelihood function fφ to this degree. 4see grove & white 2024b for details. proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 179 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the testing instrument employed in our experiments, for example, is a slider scale that records responses on the unit interval [0, 1]. a suitable likelihood would therefore be a truncated normal distribution: f(x,φ) = n (x, σ)t[0, 1] (so that f = n and φ = σ). this likelihood—which grove & white (2024a) also employ in their models of factivity—can be viewed as allowing some distribution of response errors, given the intended target response (i.e., the answer to the question). these are the ingredients we need to model the effects of the fine-grained semantics of the target and question prompt, given a particular inference task. we provide further details in section 5. 3. factive inferences. to investigate how natural language prompts modulate the distribution of participants responses, we modify an experimental paradigm developed by degen & tonhauser (2021, 2022), which they use to experimentally investigate the projective inferences triggered by factive predicates—henceforth, factive inferences. we describe the paradigm (section 3.1), then discuss previous modeling of their data within pds (section 3.2). the paradigm forms the basis for our experiments in section 4; we build directly on our previous modeling in section 5. 3.1. measuring factive inferences. degen & tonhauser’s (2021) main aim is to characterize the influence of world knowledge on factive inferences. to achieve this, they measure factive inferences in the presence of a background fact whose content they manipulate (their experiment 2b). for example, participants are given trials of the form in (2) and are asked to respond using a slider scale, with no on one end and yes on the other. (2) a. fact (which elizabeth knows): zoe is a math major. b. elizabeth asks: “does tim know that zoe calculated the tip?” c. is elizabeth certain that zoe calculated the tip? they focus on the set of twenty clause-embedding predicates listed in (3).5 (3) twenty clause-embedding predicates (degen & tonhauser 2022, p. 559, ex. 13) a. canonically factive: be annoyed, discover, know, reveal, see b. non-factive (i) non-veridical non-factive: pretend, say, suggest, think (ii) veridical non-factive: be right, demonstrate c. optionally factive: acknowledge, admit, announce, confess, confirm, establish, hear, inform, prove each embedded clause in their experiment is paired with one of two facts: either a fact intended to make the clause likely to be true (as in the example above), or a fact intended to make the clause unlikely to be true. to validate the use of these background facts for this purpose, degen & tonhauser (2021) conduct a norming experiment, in which the prior certainties about the truth of the complement clauses featured in their projection experiment are assessed independently, given 5the grouping in (3) is due to degen & tonhauser 2022 and is based on the prior literature on factivity (kiparsky & kiparsky 1970, karttunen 1971, et seq.). proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 180 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the same background facts (their experiment 2a). trials in this experiment ask participants to judge how likely the relevant clause is to be true, given one of the two background facts constructed for it. for example, participants are given trials of the form in (4) and are asked to respond using a slider scale, with impossible on one end and definitely on the other. (4) a. fact: zoe is a math major. b. how likely is it that zoe calculated the tip? degen & tonhauser find that the by-item means for the forty pairs of complement clauses and background facts, as assessed in their norming experiment, are a good predictor of the inference ratings for items featuring the same complement clauses and facts which they obtain in their experiment investigating projective inferences. 3.2. modeling factive inferences. degen & tonhauser (2022) and others (white & rawlins 2018) observe that measures of a predicate’s factivity derived from the sort of judgment data degen & tonhauser (2021) collect display gradience when the data is aggregated.6 using models developed in pds, grove & white (2024a) ask whether this apparent gradience arises due to metalinguistic uncertainty—uncertainty about whether a predicate is factive or not—or occasional uncertainty—uncertainty inherently associated with predicate meanings.7 under a metalinguistic uncertainty account, different predicates differ in the frequencies with which they trigger factive inferences across uses. under an occasional uncertainty account, predicates would license inferences with varying degrees of certainty on particular uses, similar to the manner in which a vague predicate, such as tall, can license uncertain inferences about the heights of individuals of which it is predicated. grove & white fit four models to degen & tonhauser’s data, varying whether uncertainty about either background world knowledge or factivity is encoded as metalinguistic or occasional. we extend their models here by adding an explicit model of the semantics of the natural language prompt. in degen & tonhauser’s (2021) original paradigm, this prompt is the one in (5). (5) is person certain that clause? following grove & white’s suggestion that no extant proposal posits that world knowledge should display metalinguistic uncertainty—a suggestion consistent with their modeling results—we focus specifically on two of their models: the discrete-factivity model (df), which regards uncertainty about factivity as metalinguistic and uncertainty about world knowledge as occasional, and the wholly-gradient model (wg), which regards both kinds of uncertainty as occasional. grove & white find that df performs the best in a model comparison pitting all four of their models against each other, as assessed by expected log pointwise predictive densities (elpds). 6degen & tonhauser (2022) argue that this gradience is evidence that there is no distinct classes of factive predicates. this argument is based on an apparent lack of clear clustering in the aggregate measures of different predicates’ ratings, as derived from inference judgment data collected using this and similar paradigms (white & rawlins 2018, ross & pavlick 2019). this lack of clear clustering, however, is likely a product of measurement noise: when such noise is appropriately modeled, distinct classes of factive predicates reveal themselves (kane, gantt & white 2022). 7see grove & white 2024a for the full formal details of their models. proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 181 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ they argue that this finding lends support to a view of factivity whereon it is a fundamentally discrete phenomenon, and they discuss how both conventionalist and conversationalist accounts might approach this sort of discreteness. 4. modifying the prompt. while grove & white’s results are promising, they are consistent with the possibility that the nature of degen & tonhauser’s prompt biased experimental participants toward making discrete ‘yes’ or ‘no’ judgments, even while the contribution to inference judgments made by factive predicates may be gradient. because the prompt is a polar question, and yes and no label the slider scale, participants may effectively treat their response as a binary forced choice by providing an answer near yes if they are sufficiently certain about the relevant inference, and an answer near no if they are not. if so, an a priori advantage is conferred on models regarding the contribution to inference of factive predicates as discrete and, thus, models which regard uncertainty about factive inferences as metalinguistic. to assess the effect the prompt has on participants’ responses, we conduct two experiments identical to degen & tonhauser’s, but which vary the prompt. in both, participants are provided with a degree question, which is either about the speaker’s degree of certainty (6a) or the degree of likelihood that the speaker is certain (6b). (6) a. how certain is person that clause? b. how likely is it that person is certain that clause? the prompt in (6a) was paired with a slider labeled not at all certain on the left and completely certain on the right, while the prompt in (6b) was paired with a slider labeled impossible and definitely (following degen & tonhauser’s norming experiment). the idea behind using the degree questions in (6), which involve the modal adjectives certain and/or likely, is that insofar as factivity is fundamentally gradient in nature, such degree questions should encourage participants to contact that fundamentally gradient representation, whatever it may consist in. at the very least, they should not discourage participants from relying on such a gradient representation (as a polar question might), and furthermore, they should not encourage them to discretize it. in pds, these degree questions can be seen as driving inferences based on the gradience encoded in the common ground. if factivity is fundamentally gradient in nature, factive predicates should contribute correspondingly gradient updates. all aspects of the experimental materials and methods (besides the prompts) match degen & tonhauser’s. we recruited 300 participants to give judgments for each prompt. data from 15 participants was removed in the experiment employing (6a), and data from 7 participants was removed in the experiment employing (6b); for both, we followed degen & tonhauser’s criteria. both groups of participants were recruited through amazon mechanical turk and paid $2. 5. modeling the prompt. 5.1. adding a prompt model. we fit both the df and wg models of grove & white 2024a, while manipulating the semantics of the question prompt for both.8 specifically, to model the 8we specifically use the variants of grove & white’s models that employ a truncated normal likelihood and that incorporate an anti-veridicality component. see their appendix for details. proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 182 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ prompt in (6a), we assume that the degree introduced by certain ranges over degrees of confidence rather than degrees of probability (see klecha 2012, goodman 2023), and thus that its scale is truncated relative to that of likely. we refer to this model as the confidence scale model. (7) jcertaink = λs. f ∼ λϕ.pr ( i ∼ cg(s) ϕ(i) ) ⟨λϕ, d, i.max(0,f(ϕ)−θcertain(s)) 1−θcertain(s) ≥ d, s⟩ the type of the expression in (7) is σ → p(((ι → t) → r → ι → t)×σ): it maps a proposition onto a question meaning that returns t for threshold degrees less than or equal to a degree of confidence obtained from the probability that the proposition is true in the common ground of the current state; in particular, a is certain that s is true at a threshold d if the probability of s is greater than d added to a fixed cutoff point determined by the current state, where this probability is further scaled by the size of the interval determined by the cutoff point. to model the prompt in (6b), we assign a semantics to likely according to which it introduces a degree corresponding to a probability (as opposed to a degree of confidence; see (8)), and where this degree is computed based on the corresponding semantics for certain. we refer to this model as the probability of confidence scale model. (8) jlikelyk = λs. 〈 λϕ, d, i.pr ( i′ ∼ cg(s) ϕ(i′) ) ≥ d, s 〉 note that to compute the meaning of likely [ certain s ], the degree threshold of certain needs to be saturated, in order to obtain a proposition of type ι → t which can be fed to likely. to accomplish this, we effectively assume two things: (a) that this threshold is a metalinguistic parameter; and (b) that the probability that the certainty about the prejacent s is greater than the threshold is determined by a distribution centered at the actual probability of s, which may vary by both participant and prejacent. thus we effectively attribute to each experimental participant uncertainty about the epistemic state of the propositional attitude holder denoted by the relevant embedded subject; this uncertainty, furthermore, is represented by a distribution centered at a value estimated about the participant’s background world knowledge. in the end, we effectively attribute to the bare form of certain the semantics of an (e.g., maximum standard) absolute gradable adjective. both models contrast with grove & white’s original model, which encodes a meaning for the prompt (5) according to which it asks for a value on a probability scale. we refer to this model as the probability scale model. 5.2. model fitting. we fit all models using hamiltonian monte carlo sampling as implemented in stan—specifically, cmdstanr (gabry & češnovar 2023). four chains were sampled for each model to assess convergence, with at least 6,000 warmup samples and at least 6,000 samples kept per chain. all convergence diagnostics implemented in cmdstanr were conducted. in cases where a model did not converge for reasons that can be solved by drawing more samples, the number of samples was increased until convergence was reached. even after substantial increases proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 183 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ − − − − − − − − − − experiment 1: 'how certain is x that p' experiment 2: 'how likely is it that x is certain that p' probability scale confidence scale probability of confidence scale probability scale confidence scale probability of confidence scale 1800 1900 2000 2100 1800 2000 2200 2400 semantics of prompt e lp d gradience metalinguistic occasional figure 1: elpds for the models of the data collected using the prompts (6a) and (6b). we do not show the probability of confidence scale models encoding gradience as occasional in nature, since they do not converge. error bars indicate standard errors of the elpd, as well as of the pointwise differences from the best-performing model in each facet. in the number of samples, we were unable to fit to convergence the probability of confidence scale models encoding gradience as occasional in nature. we do not show the results for this model here for this reason.9 5.3. results. we follow grove & white in reporting model comparisons using elpd. figure 1 shows the elpd for the models of the data collected using the prompts (6a) and (6b). error bars indicate standard errors of the elpd, as well as of the pointwise differences from the bestperforming model in each facet. recall that the df models are those which regard factive inferences as giving rise to metalinguistic gradience, while the wg models are those regarding it as giving rise to occasional gradience. the y-axes of each plot provide the three variants of the semantics of the prompt discussed above. across the board, the df models continue to perform the best on both datasets. meanwhile, we find that confidence scale model performs the best on the dataset containing the prompt in (6a), as expected, while the confidence scale and probability of confidence scale models perform about equally on the dataset containing the prompt in (6b) and better than grove & white’s original prompt model (the probability scale model). this result might suggest that, while we may approximate the denotation of likely well, the denotation we assign to certain is not quite correct, even if it ends up providing the correct question meaning for the prompt in (6a). we take this result to highlight the benefits of developing models in the pds framework, since it makes it straightforward to quantitatively evaluate alternative denotations, including those not considered in prior literature. 6. conclusion. our results suggest that, when analyzing data from an inference judgment task, it can be important to incorporate into one’s data analysis regime an explicit representation of the 9the fits we obtained show extremely poor performance on both datasets, though we cannot read too much into this fact, given that these models did not converge. proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 184 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ semantics of the natural language prompt used to guide participants on the task. they additionally confirm (i) that the model comparisons obtained by grove & white (2024a) do not reflect an a priori bias conferred on the discrete models by the experimental task, but rather these models’ abilities to capture the distributions of degrees of certainty associated with the inferences generated by the predicates and complement clauses tested; and (ii) that while prior work has potentially provided a good approximation to the correct semantics for predicates like certain and likely, better approximations may potentially help in studying their semantic contributions to complex inferences involving both. future research in this line will aim to leverage pds to explore the space of possible denotations for such predicates via explicit model comparison of the form we employ here. beyond improving our understanding of the semantics of predicates like likely, we believe that the approach we have taken here can help us to better understand the fine-grained semantics of attitude predicates like certain, as well—a point supported by our success in modeling (6a). one clear opportunity for improving the denotation of predicates like certain is to investigate whether or not our current denotation takes the attitude holder’s epistemic state into account in the right way: here we assume that the relevant contextual standard is fixed, while degrees of certainty vary according to a distribution determined by a given participant’s background knowledge. further investigating how an attitude holder’s epistemic state should be incorporated into the semantics of certain is thus a potentially interesting future direction that is imminently feasible within pds. references cooper, robin et al. 1996. using the framework. technical report lre 62-051 d-16. the fracas consortium. dagan, ido, oren glickman & bernardo magnini. 2006. the pascal recognising textual entailment challenge. en. in joaquin quiñonero-candela et al. (eds.), machine learning challenges. evaluating predictive uncertainty, visual object classification, and recognising tectual entailment, 177–190. berlin, heidelberg: springer. https://doi.org/10.1007/ 11736790_9. degen, judith & judith tonhauser. 2021. prior beliefs modulate projection. open mind 5. 59–70. https://doi.org/10.1162/opmi_a_00042. https://doi.org/10.1162/ opmi_a_00042 (25 april, 2023). degen, judith & judith tonhauser. 2022. are there factive predicates? an empirical investigation. en. language 98(3). publisher: linguistic society of america, 552–591. https://doi. org/10.1353/lan.0.0271. https://muse.jhu.edu/article/864635 (28 november, 2022). farkas, donka f. & kim b. bruce. 2010. on reacting to assertions and polar questions. journal of semantics 27(1). 81–118. https://doi.org/10.1093/jos/ffp010. https: //doi.org/10.1093/jos/ffp010 (28 april, 2024). frank, michael c. & noah d. goodman. 2012. predicting pragmatic reasoning in language games. science 336(6084). publisher: american association for the advancement of science, 998–998. https://doi.org/10.1126/science.1218633. https:// www.science.org/doi/10.1126/science.1218633 (20 june, 2022). proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 185 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ gabry, jonah & rok češnovar. 2023. cmdstanr. tech. rep. https://mcstan.org/ cmdstanr/index.html. gantt, william, benjamin kane & aaron steven white. 2020. natural language inference with mixed effects. arxiv:2010.10501 [cs]. https://doi.org/10.48550/arxiv.2010. 10501. http://arxiv.org/abs/2010.10501 (17 june, 2024). ginzburg, jonathan. 1996. dynamics and the semantics of dialogue. in jerry seligman & dag westerståhl (eds.), logic, language, and computation, vol. 1, 221–237. stanford: csli publications. goodman, jeremy. 2023. degrees of confidence are not subjective probabilities. in proceedings of sinn und bedeutung 28. goodman, noah d. & michael c. frank. 2016. pragmatic language interpretation as probabilistic inference. en. trends in cognitive sciences 20(11). 818–829. https://doi.org/10. 1016/j.tics.2016.08.005. https://www.sciencedirect.com/science/ article/pii/s136466131630122x (11 february, 2021). groenendijk, jeroen & martin stokhof. 1984. studies on the semantics of questions and the pragmatics of answers. amsterdam: university of amsterdam dissertation. https://stokhof. org/wp-content/uploads/2020/09/groenendijk-stokhof_ssqpa.pdf. grove, julian & jean-philippe bernardy. 2023. probabilistic compositional semantics, purely. en. in katsutoshi yada et al. (eds.), new frontiers in artificial intelligence (lecture notes in computer science), 242–256. cham: springer nature switzerland. https://doi.org/ 10.1007/978-3-031-36190-6_17. grove, julian & aaron steven white. 2024a. factivity, presupposition projection, and the role of discrete knowlege in gradient inference judgments. lingbuzz published in: submitted. https://ling.auf.net/lingbuzz/007450 (25 april, 2024). grove, julian & aaron steven white. 2024b. probabilistic dynamic semantics. lingbuzz published in: https://ling.auf.net/lingbuzz/008478 (20 october, 2024). hausser, roland & dietmar zaefferer. 1978. questions and answers in a context-dependent montague grammar. en. in f. guenthner & s. j. schmidt (eds.), formal semantics and pragmatics for natural languages, 339–358. dordrecht: springer netherlands. https://doi.org/ 10.1007/978-94-009-9775-2_12. https://doi.org/10.1007/978-94009-9775-2_12 (19 april, 2024). hausser, roland r. 1983. the syntax and semantics of english mood. en. in ferenc kiefer (ed.), questions and answers, 97–158. dordrecht: springer netherlands. https://doi.org/ 10.1007/978-94-009-7016-8_6. https://doi.org/10.1007/978-94009-7016-8_6 (19 april, 2024). jasbi, masoud, brandon waldon & judith degen. 2019. linking hypothesis and number of response options modulate inferred scalar implicature rate. english. frontiers in psychology 10. publisher: frontiers. https://doi.org/10.3389/fpsyg.2019.00189. https://www.frontiersin.org/journals/psychology/articles/10. 3389/fpsyg.2019.00189/full (3 july, 2024). kane, benjamin, will gantt & aaron steven white. 2022. intensional gaps: relating veridicality, factivity, doxasticity, bouleticity, and neg-raising. en. semantics and linguistic theory proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 186 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 31(0). 570–605. https://doi.org/10.3765/salt.v31i0.5137. https: //journals.linguisticsociety.org/proceedings/index.php/salt/ article/view/31.029 (11 may, 2023). karttunen, lauri. 1971. some observations on factivity. paper in linguistics 4(1). publisher: routledge _eprint: https://doi.org/10.1080/08351817109370248, 55–69. https://doi.org/ 10.1080/08351817109370248. https://doi.org/10.1080/08351817109370248 (26 june, 2023). karttunen, lauri. 1977. syntax and semantics of questions. linguistics and philosophy 1(1). publisher: springer, 3–44. https://www.jstor.org/stable/25000027 (17 june, 2024). kiparsky, paul & carol kiparsky. 1970. fact. en. in progress in linguistics, 143–173. de gruyter mouton. https://doi.org/10.1515/9783111350219.143 (9 june, 2023). klecha, peter. 2012. positive and conditional semantics for gradable modals. en. proceedings of sinn und bedeutung 16(2). number: 2, 363–376. https://ojs.ub.uni-konstanz. de/sub/index.php/sub/article/view/433 (14 december, 2023). maccartney, bill. 2009. natural language inference. en. stanford: stanford university dissertation. https://www-nlp.stanford.edu/~wcmac/papers/nli-diss.pdf. roberts, craige. 2012. information structure: towards an integrated formal theory of pragmatics. en. semantics and pragmatics 5. 6:1–69. https://doi.org/10.3765/sp.5.6. https://semprag.org/index.php/sp/article/view/sp.5.6 (26 july, 2023). ross, alexis & ellie pavlick. 2019. how well do nli models capture verb veridicality? in proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (emnlp-ijcnlp), 2230–2240. hong kong, china: association for computational linguistics. https://doi. org/10.18653/v1/d19-1228. https://aclanthology.org/d19-1228 (26 june, 2023). stalnaker, robert. 1978. assertion. in peter cole (ed.), pragmatics, vol. 9, 315–332. new york: academic press. white, aaron steven & kyle rawlins. 2018. the role of veridicality and factivity in clause selection. in sherry hucklebridge & max nelson (eds.), nels 48: proceedings of the forty-eighth annual meeting of the north east linguistic society, vol. 48, 221–234. university of iceland: glsa (graduate linguistics student association), department of linguistics, university of massachusetts. xiang, yimei. 2021. a hybrid categorial approach to question composition. en. linguistics and philosophy 44(3). 587–647. https://doi.org/10.1007/s10988-020-09294-8. https://doi.org/10.1007/s10988-020-09294-8 (26 september, 2024). proceedings of elm 3: 176-187, 2025 julian grove and aaron steven white: modeling the prompt in inference judgment tasks. 187 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction kurt erbach, magnus poppe, & cornelia ebert∗ abstract. we report results of experiments on pronoun and presupposition binding across modalities. we show that ordinary pronouns (in the spoken/written domain) can be dynamically bound to gesturally introduced discourse referents and that presuppositions induced by spoken/written triggers (via e.g. ‘again’ or ‘too’) can be bound likewise. these experiments support research that has proposed the existence of cross-modal binding to motivate a formal framework that can account for interaction of various input of linguistic content from different dimensions and modalities. keywords. multimodal; gesture; anaphora; pronouns; presupposition 1. introduction. ebert et al. (2020) (based on ebert & ebert 2014) suggest a formal framework for gesture semantics where certain iconic and pointing co-speech gestures introduce discourse referents that can serve as antecedents in anaphoric reference. crucially, this necessitates a unidimensional dynamic system that allows for binding effects across dimensions and, in this case, modalities. ebert et al. (2020) argue that co-speech gestures contribute propositional non-at-issue information by default. the further argue that this information arises as ‘constructional’ meaning due to the co-occurrence of gesture and speech. in their dynamic semantic framework, gestures introduce discourse referents for rigid designators as their core ‘lexical’ meaning: when a pointing gesture is performed this triggers the introduction of a discourse referent that is identified with the rigid concept of the gesture referent, and can be anaphorically picked up across dimensions—e.g. by a pronoun. while the introduction of discourse referents by gesture has been claimed and implemented in the formal system of ebert et al. (2020), this has not been experimentally demonstrated. here we show that dynamic binding across dimensions can be made with respect to both pronouns and presupposition triggers like again. in the constructed example (1-a), where both cake and cookies are contextually salient, the pointing co-speech gesture, extending an index-finger towards the cake, is assumed to introduce a discourse referent for the gesture concept, resulting in an interpretation somewhat like ‘have you eaten (some) cake?’. the discourse referent for the gesture concept is assumed to be able to bind to pronouns like it in the hypothetical response (1-b). on the other hand, if (1-a) had included a hand-over-stomach gesture to indicate being full (1-c) and crucially not introducing some food as a discourse referent, then presumably responding up with (1-b) would be infelicitous as it is unclear what it is supposed to be bound to. moreover, if the question with the pointed-to referent (1-a) were responded to in such a way that answers the question albeit ignoring the referent (1-d), then this too was predicted to be marked under the assumption that the gesture referent is introduced, but not picked up. yet, one would expect this case to be not as bad as an utterance with a pronoun that cannot be resolved to any referent. ∗we would like to thank the organizers, reviewers, and audiences of elm3, linguistic evidence 2024, and xpragit.2024 for their feedback on this project and the opportunity to present our work. authors: kurt erbach, goethe-university frankfurt, saarland university (erbach@hhu.de); magnus poppe, goethe-university frankfurt; cornelia ebert, goethe-university frankfurt (ebert@lingua.uni-frankfurt.de). proceedings of elm 3: 152-162, 2025 c©2025 kurt erbach, magnus poppe, and cornelia ebert published by the lsa with permission of the author(s) under a cc by license. 152 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ (1) a. have you eatenpointing to cake? b. it was too sweet for me. c. have you eatenhand over stomach? d. yeah, a few too many cookies. similarly, the jogging gesture in (2-a) is assumed to add the propositional content that paul was jogging (when the speaker met him). if (2-b) is said in response to (2-a), it is assumed that the presupposition that is triggered by again in (2-b), namely that paul went jogging before, is satisfied by the the propositional information about jogging given in the visual modality in (2-a). if (2-b) was a response to (2-c), the absence of a jogging gesture would mean the presupposition triggered by again would not be satisfied, at least under the assumption that people don’t commonly meet while jogging and hence such a proposition cannot be accommodated. conversely, a follow up like (2-d) is presumably odd following a jogging gesture (2-a) under the assumption that people do not jog in cafés, but following (2-c) ought to be fine assuming people often meet in cafés. (2) a. yesterday i met pauljogging gesture? b. he went jogging again today. c. yesterday i met paulpointing behind self? d. was it in the café again? the paradigms in (1) and (2) exemplify data that are used in our experiments by which we seek to support the idea that gesture can introduce discourse referents that can be picked up by anaphora. in this paper, we describe a series of experiments that test this idea via these paradigms and that ultimately show the sorts of acceptability contrasts that are predicted. in section 2, we describe the design of our experiments, in section 3 we report the results, and in section 4 we describe a followup experiment and its results showing that the acceptability of pronoun binding with pointed-to referents depends on grammatical gender agreement. in section 5, we discuss the results of the experiments, namely that these experiments can be taken as evidence of dynamic binding, wherein verbal anaphora can be bound to content introduced by gestures. 2. experiment design. two experiments were designed in german to test the contrasts demonstrated in (1) and (2). the complementary designs allowed each to be used as filler for the other. both experiments had two factors each with two levels—gesture (pointing/non-iconic or iconic/non-pointing) and pronoun or presupposition (present or absent)—yielding four treatments, each of which was predicted to be judged as either felicitous (accept) or infelicitous (reject). the four conditions for the respective pronoun and presupposition binding experiments are listed in tables 1 and 2. gesture pronoun prediction examples pointing present felicitous (1-a)+(1-b) pointing absent infelicitous (1-a)+(1-d) non-pointing present infelicitous (1-c)+(1-b) non-pointing absent felicitous (1-c)+(1-d) table 1: conditions for pronoun-binding experiment 10 minimal pairs resembling (1) and (2) were distributed across lists for a 2x2 latin square proceedings of elm 3: 152-162, 2025 kurt erbach, magnus poppe, and cornelia ebert: experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction. 153 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ gesture presupposition factor level examples iconic present felicitous (2-a)+(2-b) iconic absent infelicitous (2-a)+(2-d) non-iconic present infelicitous (2-c)+(2-b) non-iconic absent felicitous (2-c)+(2-d) table 2: conditions for presupposition-binding experiment design (i.e. 10 items per list), and we recruited 80 native german speaking participants via prolific (20 per list), so there were 200 responses per condition1. in a variation of the covered-box task (cf. fanselow et al. 2022), the sentence pairs were presented with the context, e.g. (1-c) in video form, and the follow-up, e.g. (1-b), being in written form as one choice in a pair of alternatives, the other being ‘covered’ (lit. “[geschwärzt]” (‘redacted’)). participants were instructed that only one of the alternatives was a reasonable follow-up to the context, and they should select the reasonable response. all experimental materials can be found in the repository at https://osf.io/3hvyt/, along with the results and r code used for analysis. prior to performing the actual experiment, each participant performed four practice trials to get familiar with the design. the first two practice trials included instructions but no requirement to select the answer as predicted. the latter two practice trials included no instructions. the practice items followed the same basic formula albeit not with respect to pronoun reference and with different gestures and sentences in each utterance. to control for quality, four attention-check questions were added using the same videos as the practice trials. underneath the video was a statement that this question is designed to check that the participant is not a bot—following silber et al. (2022), who showed that such statements increase accuracy in attention check questions—and the question asked something about the appearance of the speaker in the video—e.g. ‘is the speaker wearing glasses?’. answers from the participants that got more than one attention-check question wrong were discarded. the experiments were preregistered on osf at https://osf.io/35jbw. 3. results. one participant each in groups 1 and 3 failed more than one attention check test as did two participants in group 4. these participants’ answers were all removed from the results as well as those of randomly selected participants in groups 1, 2, and 3 such that each group had the same number of participants remaining (n=18). therefore, 180 responses were analyzed per condition (cf. the originally collected 200 responses per condition). 3.1. pronoun binding experiment results. the results for the pronoun experiment are summarized in figure 1. in this experiment, follow-up sentences with pronouns to be bound to the discourse referent introduced with a pointing gesture (1-a)+(1-b) were largely accepted (n=115), and, surprisingly, those without such a pronoun (1-a)+(1-d) were accepted nearly as much (n=105). with non-pointing gesture contexts, follow-ups with pronouns to bind (1-c)+(1-b) were not accepted (n=63) unlike those with non-pointing continuations ((1-c)+(1-d), n=133). 1because there were a total of 10 items per condition, there was an uneven distribution of conditions across lists— e.g. three pointing–present items in list 2, two pointing–present items in list 3, two non-pointing–present items in list 2, three non-pointing–present items in list 3, etc. proceedings of elm 3: 152-162, 2025 kurt erbach, magnus poppe, and cornelia ebert: experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction. 154 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 1: number of follow-up acceptances in pronoun experiment; max-n = 180 to understand whether or not there is an effect of pointing in the resolution of pronoun binding with otherwise unspecified discourse referents—i.e. to understand if there is a difference between (1-a)+(1-b) and (1-c)+(1-b)—the results were analyzed with a mixed-effects model with binomial error distribution constructed in r (r core team 2015) using the lmertest package (kuznetsova et al. 2019). the maximal random effects structure justified by our experimental design should include a by-participant random effect, but this model did not converge, so we simplified the random effects structure by not allowing correlation between the random slope and random intercept for subjects (barr et al. 2013). the full model translates to: response & bind * gesture + (1|questionid). the summary of the model is in table 3; the main effect of bind–gesture interaction is statistically significant (p = 0.0026). to summarize, there is a very low likelihood that we would see these fixed effects estimate se z p intercept 1.3091 0.3988 3.283 0.00103 bind -2.0635 0.5567 -3.707 0.00021 gesture -0.8904 0.5516 -1.614 0.10646 bind:gesture 2.3420 0.7778 3.011 0.00260 random effects name variance std.dev. questionid (intercept) 1.172 1.083 table 3: summary of glmer model for pronoun experiment results if there was no interaction between binding pronouns and gesture. we take this as support of the proposal that discourse referents can be introduced with pointing gestures, and that pronouns proceedings of elm 3: 152-162, 2025 kurt erbach, magnus poppe, and cornelia ebert: experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction. 155 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ can then be bound to these discourse referents. and crucially, if no pointing occurred during an utterance in a context where possible discourse referents are not otherwise clear, then a follow-up with a to-be-bound pronoun is likely to be rejected. 3.2. presupposition binding experiment results. the results for the presupposition experiment are summarized in figure 2. in this experiment, follow-up sentences with presupposifigure 2: number of follow-up acceptances in presupposition experiment; max-n = 180 tion triggers like again to be bound to the propositional referent introduced with an iconic gesture introducing some propositional content (2-a)+(2-b) were largely accepted (n=129) while the same follow-ups were rejected (n=72) if the gesture did not introduce such a propositional referent (2-c)+(2-b). our control items were also judged as predicted: when iconic gesture contexts were followed with a statement that did not refer to the gesture, rather it contained an item that clashed with the gesture, (2-a)+(2-d), these follow-ups were rejected (n=86), but when a context gesture did not introduce a discourse referent (2-c)+(2-d), the same follow-ups were accepted (n=138). to understand whether or not there is an effect of iconic gestures in the resolution of presupposition trigger binding with otherwise unspecified discourse referents—i.e. to understand if there is a difference between (2-a)+(2-b) and (2-c)+(2-b)—the results were analyzed with a mixed-effects model with binomial error distribution constructed in r (r core team 2015) using the lmertest package (kuznetsova et al. 2019). the maximal random effects structure justified by our experimental design translates to: response & bind * gesture + (1|questionid) + (1|participant). the summary of the model is in table 4; the main effect of bind–gesture interaction is statistically significant (p = 0.00016). to summarize, there is a very low likelihood that we would see these results if there was no interaction between binding presupposition triggers and to propositional referents introduced with an iconic gesture. we take this to support the claim that propositional proceedings of elm 3: 152-162, 2025 kurt erbach, magnus poppe, and cornelia ebert: experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction. 156 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ fixed effects estimate se z p intercept -0.3056 0.4885 -0.626 0.53152 bind 1.4737 0.6809 2.164 0.03045 gesture 1.8712 0.6891 2.715 0.00662 bind:gesture -3.6653 0.9710 -3.775 0.00016 random effects name variance std.dev. participant (intercept) 0.353 0.5941 questionid (intercept) 1.912 1.3829 table 4: summary of glmer model for presupposition experiment discourse referents can be introduced with iconic gestures, and that presupposition triggers can then be bound to these discourse referents. moreover, if no iconic gesture occurred during an utterance in a context where possible discourse referents are not otherwise clear, then a follow-up with a to-be-bound presupposition trigger is likely to be rejected. 3.3. interim discussion. in general, the results can be interpreted as showing that participants accept follow-ups with pronouns and presupposition triggers that can be bound to discourse items introduced with pointing and iconic gestures respectively more readily than when there are no possible binders for these pronouns and presuppositions. in both experiments these items were accepted more than 50% (n=90) of the time: pronouns (n=115); presupposition triggers (n=129). it may be worth noting that these rates of acceptance are closer to random acceptance than complete acceptance (100%, n=180), which may be due to a number of reasons: first, the entire experimental paradigm is a bit unnatural in that the contexts were presented in video format and the follow-ups were presented not only in a written format, but also in a covered box task, both of which may have led to fewer than expected acceptances of follow-up items. this possibility might be supported by the fact that even our “acceptable” contrast items (i.e. non-pointing (1-c) + item-to-bind-absent (1-d) for the pronoun experiment, and non-iconic (2-c) + item-to-bind-absent (2-d) for the presupposition trigger experiment) had acceptance rates closer to random choice than complete acceptance: n=133 in the pronoun experiment and n=135 in the presupposition trigger experiment. another reason for low acceptance might be the covered-box paradigm, where participants were invited to imagine that one of the two follow-ups was definitely acceptable and the other definitely unacceptable, and only one of the two was shown. this may have contributed to the given acceptance rates because participants may have imagined even more acceptable responses than the ones given in certain cases. another notable result is the fact that, in the pronoun experiment, follow-up items that ignored the discourse referent introduced by pointing, (1-a)+(1-d), were accepted more than half the time (n=105, 58.33%). while this was the closest count to random acceptance (50%), two of the other three conditions predicted to be marked were similarly close to random acceptance: for gesture-iconic + presupposition-absent, (2-a) + (2-d), follow-ups were rejected 47.77% of the time (n=86), and for the condition gesture-non-iconic+presupposition-present, (2-c)+(2-b), follow-ups were rejected 40% of the time (n=72). this too could speak to the issues with the experimental design discussed above. a final alternative may be that, as mentioned above, it simply proceedings of elm 3: 152-162, 2025 kurt erbach, magnus poppe, and cornelia ebert: experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction. 157 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ might be natural to ignore a discourse referent introduced with a co-speech pointing gesture in the scope of a question, and for sure not as bad as having to interpret a to-be bound pronoun that has no antecedent. 4. pronoun binding follow-up experiment. we conducted a follow-up to the pronoun experiment to test for an effect of pronoun agreement when discourse pronouns in follow-up sentences are meant to bind to discourse referents that are not lexically specified but possibly pointed to with a pointing gesture. the data from the previous pronoun binding study consisted of 10 items testing pronoun resolution and gesture, 8 of which contained pronouns that agreed only with the unspoken noun that could refer to the object pointed to as opposed to also agreeing with the other item visible in the context. for example, with the question “hast du schon gegessen?” (‘have you already eaten’) a piece of cake was pointed to during ‘already eaten’; note a package of cookies was also visible in this context. the participants were asked to choose between a pair of responses, one visible, one not, and in the pronoun-present condition, the visible response was “er war zu süß für mich.” (‘it.m was too sweet for me’). crucially, the german noun kuchen (‘cake.m’) has masculine gender which could be used to bind the pronoun in the follow-up question, while the alternative object, which was not pointed to, a package of cookies (kekse ‘cookies’) did not agree in gender with the pronoun in the visible response. because such a bias existed in 80% of the target items, if pronoun agreement with an unspoken noun was sufficient without pointing, then we might have expected that the visible responses were selected for that 80% of control items. instead, what we saw is that visible responses were selected for only 35% (n=63) of the control items where a pronoun agreed in gender with one of the visible items that were not pointed to. in contrast, when pointing was included in the context, the visible responses were selected 63.89% (n=115) of the time. while the inferential statistics for the previous study found that no interaction between gesture and pronoun-to-bind is unlikey (p = 0.0026), strictly speaking, the interaction of gesture and pronoun agreement has not been tested. in other words, while it seems that the presence of an object that, when named, would agree with the pronoun in question is not sufficient to accept statements that contain pronouns meant to bind with such objects, it is unclear if it is the pointing gesture alone or gesture+agreement interaction that drives the acceptance of follow-ups in the gesture+pronoun-to-bind condition. by looking at pronoun agreement and gesture interaction, the present study provides the opportunity to see how participants behave when the visible responses contain a pronoun that is unlikely to be bound to an unspoken noun referring to one of the objects. in other words, one of the conditions in the covered box paradigm will contain visible responses that do not agree in gender with the visible objects in the context that can be referred to with nouns that could serve as an antecedent for said pronoun. by minimally modifying the design of the previous study, the following paradigm (3) and conditions (table 5) were developed to test the interaction of gesture and agreement. (3) a. hast du schon gegessen? (‘have you eatenpointing to cake.m?’) b. ja. der war aber zu süß für mich. (‘yes, but itdpro.m.sg was too sweet for me’). c. hast du schon gegessen? (‘have you eatenhand over stomach?’) d. ja. die war aber zu süß für mich. (‘yes, but itdpro.f.sg was too sweet for me’). proceedings of elm 3: 152-162, 2025 kurt erbach, magnus poppe, and cornelia ebert: experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction. 158 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ the first condition in table 5 is the same as the previous pronoun binding experiment though we switched to using d-pronouns because these have been shown to be preferable in spoken modality (patil et al. 2023). the second condition tests whether a pronoun that does not agree with the pointed-to item is acceptable: in (3-d), the pronoun die has feminine grammatical gender and therefore does not agree with kuchen (‘cake.m’), which is the unspoken noun denoting the object pointed to in (3-a). the third condition is again the same as in the previous experiment albeit with the d-pronoun, and the fourth condition completes the set by testing to see whether a non-agreeing pronoun is acceptable when no specific item has been mentioned. according to the results of our previous study, this should also be rejected as there is no item mentioned to agree with. apart gesture agreement factor level examples pointing agree felicitous (3-a)+(3-b) pointing disagree infelicitous (3-a)+(3-d) non-pointing agree infelicitous (3-c)+(3-b) non-pointing disagree infelicitous (3-c)+(3-d) table 5: conditions for pronoun-agreement experiment from these modifications, the follow-up experiment followed the same design as procedure as the previous experiments. 4.1. results of pronoun agreement experiment. all 80 participants answered at least 75% of the control items correctly so no answers had to be thrown out. the results for the pronoun agreement experiment are summarized in figure 3. in relative terms, the results of this experiment figure 3: number of follow-up acceptances in pronoun follow-up experiment; max-n = 200 followed the predictions in that the pointing–agree condition, (3-a)+(3-b), received the highest number of acceptances (n=154; 77%) making it clearly acceptable, as predicted, while all others were markedly less acceptable. in particular, the pointing–disagree condition, (3-a)+(3-d), received an above chance (50%) number of accepts (n=114; 57%), which we took to be acceptable in the previous experiments, and therefore deviates from our predictions in these terms. the non-pointing conditions both received a lower than chance number of accepts: non-pointing–agree n=94 (47%) and non-pointing–disagree n=72 (36%). proceedings of elm 3: 152-162, 2025 kurt erbach, magnus poppe, and cornelia ebert: experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction. 159 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ to analyze the effects of gesture and agreement on the acceptability of the follow-up items, the results were analyzed with a mixed-effects model as before. the maximal random effects structure justified by our experimental design translates to the following: response & agreement * gesture + (1|participant) + (1|questionid). the results of the model are summarized in table 6. as seen fixed effects estimate se z p (intercept) -0.7860 0.3769 -2.085 0.037 agreement 0.6206 0.5160 1.203 0.229 gesture 1.1294 0.5139 2.198 0.028 agreement:gesture 0.5603 0.7297 0.768 0.443 random effects name variance std.dev. participant (intercept) 0.5369 0.7327 questionid (intercept) 1.0314 1.0156 table 6: summary of glmer model for pronoun follow-up experiment in table 6, there is a main effect of gesture (p=0.028) but not of agreement (p=0.229) or gesture– agreement interaction (p=0.443). an anova between this model against one without interaction between agreement and gesture does not show a significant difference (p=0.4454), however an anova between the model without interaction between agreement and gesture and one only with gesture as a fixed effect—i.e. without agreement as a fixed effect at all—does show an effect of agreement (p=0.0182). 4.2. discussion of agreement experiment. we take the results of the anova between linear models with and without agreement as evidence that there is an effect of agreement on the acceptance of follow-ups in the pronoun agreement experiment. in other words, while participants did accept pronouns at a rate above chance when they do not agree in gender with the unspoken names of objects visible in the context, such pronouns are considered marked to those that do agree. given that the linear model shows no effect of agreement (in contrast to the anova) despite showing an effect of gesture, one might also be inclined to assume that the pointing gesture drives the agreement in these. this might be corroborated by the above-chance rate of acceptance in the pointing–disagree condition. looking back at our stimuli, they generally also did not agree with the unspoken noun that would denote the object not pointed to, though it could have for a few of the items depending on exactly how the not-pointed-to object was interpreted. it also could have been the case that interpretations we did non anticipate were thought of by the participant—e.g. “die torte” (‘thef layer-cake’) instead of “der kuchen” (‘them cake’). in short, though it is clear that the participants strongly preferred pronouns that agree with unspoken nouns denoting pointed-to objects to those that do not agree, further work is necessary to understand the acceptance rate of pronouns in the pointing–disagree condition. 5. discussion. the results of the three experiments support the proposals by ebert & ebert (2014), ebert et al. (2020), namely that gestures can introduce discourse referents that pronouns and presupposition triggers can bind to. in the first experiment, we showed that participants largely accepted the idea of using pronouns to refer to discourse referents introduced with pointing gestures, proceedings of elm 3: 152-162, 2025 kurt erbach, magnus poppe, and cornelia ebert: experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction. 160 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ in contrast to instances when pronouns were used to refer to objects visible in the context but were not pointed to. in the second experiment, we showed that participants largely accepted the idea of using presupposition triggers like again to refer to discourse referents introduced with iconic gestures, in contrast to instances when these referents were ignored, or when presupposition triggers did not have a clear referent to bind to. finally, in the third experiment we showed that participants were more likely to accept pronouns that agree in gender with the unspoken noun of a pointed-to object than pronouns that disagree. in other words, we have seen the acceptability of anaphora binding to discourse referents introduced with gender in three distinct experiments. within and between the two pronoun experiments, we also saw a contrast between instances where pronouns are used without clear discourse referents in the preceding context and instances where they are. in short, the question is, why does pointing play such a key role when pronoun gender can otherwise disambiguate what is being referred to? recall that, following pointing gestures, the referring to the pointed-to-but-unnamed object with a pronoun that agrees in gender with potential noun is accepted at a relatively high rate (63% in the first experiment and 77% in the follow-up). meanwhile, the same sentence is generally rejected if the object is not pointed to in the preceding context (35% acceptance in the first experiment and 47% in the follow-up). given the effect of agreement in the follow-up, why do participants reject pronouns at these rates? and at the same time, why do they accept pronouns that do not agree with nouns denoting discourse referents introduced by pointing at a relatively high rate (57%)? we argue that the account of the semantics of gesture argued for in ebert et al. (2020) (based on ebert & ebert 2014) can straightforwardly account for the acceptance rates of pronouns seen in these experiments. in particular, what is argued is that co-speech gestures (which were used in the above experiments) are interpreted as follows (4): the spoken lexical items introduce a predicate np(x), where items subscripted by p are at-issue content and those subscripted by p* are not-atissue content (see anderbois et al. 2015); the gesture introduces a discourse referent, z, which is equivalent to a rigid designator, ig, which, applied to all worlds, yields the same designated gesture referent g; and the temporal alignment between the lexical items and the gesture introduces a similarity predicate, simp∗(x, z), that applies to the argument of the lexical predicate, i.e. the concept introduced in speech, and the rigid gesture concept, and requires the two—what is introduced in speech and what is introduced via gesture—to be similar in certain contextually relevant respects. (4) [x] ∧np(x) ∧ [z] ∧ z = ig ∧ simp∗(x, z) what this means for the results of our study is that, while pointing introduces an object z that can be picked up in later discourse, exactly how this object is named in speech still remains unclear. the somewhat low rate of acceptance of follow-ups with pronouns agreeing with unspoken nouns denoting the pointed-to objects (n=154; best possible 200) fits with this analysis given some bridging must be assumed: since the pronoun is assumed to bind to the referent z that has been introduced by the gesture, the participants still have to find a spoken language expression for this referent. but there are still some uncertainties involved at this point—participants have to bridge the pronoun’s grammatical gender to some noun concept, which they have to provide themselves. in other words, the difference between accepting der (‘the.m’), versus die (‘the.f’), when the speaker in the video points to cake depends on whether or not participants decide to name and refer to the object pointed to as kuchen (‘cake.m’). proceedings of elm 3: 152-162, 2025 kurt erbach, magnus poppe, and cornelia ebert: experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction. 161 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ this analysis also leaves room for pronouns that do not agree with the unspoken noun that denotes the object in question to nevertheless bind to these discourse referents. at the same time, while non-pointing gestures can be interpreted in the same way, the objects they would point to would differ from those referred to by the pronoun. for example, if the non-pointing gesture in (1-c) “points” to a state of being full or having eaten, then the pronoun in (1-b) could not bind to this object because a state of being full cannot be the argument of ‘it was too sweet for me.’ in other words, because the context did not provide any salient discourse referents for pronouns in items like (1-b), participants accepted these items less. in summary, we have provided experimental evidence that supports the formal model in ebert et al. (2020) in which gestures introduce discourse referents that pronouns and presupposition triggers like ‘again’ can bind to. we have also demonstrated the effect of gender on binding pronouns to discourse referents introduced by a pointing gesture. lastly, we argued that the semantic analysis in ebert et al. (2020) can account for the fact that pronouns that do not agree with nouns denoting pointed-to objects are nevertheless generally accepted. references anderbois, scott, adrian brasoveanu & robert henderson. 2015. at-issue proposals and appositive impositions in discourse. journal of semantics 32(1). 93–138. barr, dale j, roger levy, christoph scheepers & harry j tily. 2013. random effects structure for confirmatory hypothesis testing: keep it maximal. journal of memory and language 68(3). 255–278. ebert, christian, cornelia ebert & robin hörnig. 2020. demonstratives as dimension shifters. in proceedings of sinn und bedeutung, vol. 24 1, 161–178. ebert, cornelia & christian ebert. 2014. gestures, demonstratives, and the attributive/referential distinction. handout of a talk given at semantics and philosophy in europe (spe 7), berlin 28. fanselow, gisbert, malte zimmermann & mareike philipp. 2022. accessing the availability of inverse scope in german in the covered box paradigm. glossa: a journal of general linguistics 7(1). kuznetsova, alexandra, per b. brockhoff & rune h. b. christensen. 2019. ‘lmertest’ r package version 3.1-0. https://cran.r-project.org/web/packages/lmertest/ index. patil, umesh, stefan hinterwimmer & petra b schumacher. 2023. effect of evaluative expressions on two types of demonstrative pronouns in german. glossa: a journal of general linguistics 8(1). r core team. 2015. r: a language and environment for statistical computing. foundation for statistical computing. http://www.r-project.org/. silber, henning, joss roßmann & tobias gummer. 2022. the issue of noncompliance in attention check questions: false positives in instructed response items. field methods 34(4). 346–360. proceedings of elm 3: 152-162, 2025 kurt erbach, magnus poppe, and cornelia ebert: experimental findings for a cross-modal account of dynamic binding in gesture-speech interaction. 162 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ cause, make, and force as graded causatives angela cao, aaron steven white, & daniel lassiter* abstract. we investigate the semantics of the causal verbs cause, make, and force as used in the construction x {caused/made/forced} y (to) z. the predominant approach to analyzing verbs of causing has been to argue that they convey some version of sufficiency, but it has also been suggested that intention or possible alternatives may also factor into the semantics of the verbs. using sequences of tic-tac-toe states as experimental stimuli, we measure the three possible contributing factors in each stimuli and ask participants whether each verb is appropriate for describing the sequence. we find experimental support for a differentiating semantics of these verbs, in which no single predictor is the sole factor in when each verb is appropriate. keywords. semantics; causatives; causal models; psycholinguistics 1. introduction. the predominant approach to analyzing verbs of causing has been to argue that they convey some version of sufficiency, which is measured given parameters of a causal situation (nadathur & lauer 2020, lauer & nadathur 2018, glass 2023, schulz 2011). notably, some of this work has leveraged structural causal models (scms; pearl, 2009) to model and make predictions about how we use these verbs (baglini & siegal 2021, nadathur & siegal 2022, schulz 2011). in this paper, we argue that the semantics of causing verbs encode not only sufficiency but also intention and the number of feasible alternative actions. to support this argument, we provide experimental evidence for a differentiating semantics of three causing verbs using explicitlydefined causal models, which enable us to calculate measures derived from the stimuli. we are thus able to quantify concepts including sufficiency and use them as predictors of judgements. our objects of study are the three english periphrastic causatives cause, make, and force. we investigate the meaning of these three verbs when used in the linguistic constructions of the form in (1). (1) x    caused made forced    y (to) z. where x and y refer to entities that can be agents, and z describes an action—e.g. (2). (2) [the pirate]x forced [the prisoner]y to [walk down the plank]z . these verbs are of interest because they are clearly not interchangeable, in spite of the fact that they all seem to express a similar kind of causation. consider the following examples, which were identified via pre-existing datasets (cao et al. 2022, williamson et al. 2023, davies 2008–) and modified (where [] indicates insertion/deletion/replacement) for our purposes: (3) a. a cat [...] caused himself [to] look as much as possible like a doctor... *this research underwent ethical review by the lel research ethics panel at the university of edinburgh. we would like to thank lelia glass, julian grove, joshua knobe, and the participants at elm for useful feedback. corresponding author: angela cao, university of rochester (acao9@ur.rochester.edu); also, aaron steven white, university of rochester and daniel lassiter, university of edinburgh. proceedings of elm 3: 88-103, 2025 c©2025 angela cao, aaron steven white, and daniel lassiter published by the lsa with permission of the author(s) under a cc by license. 88 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ b. a cat [...] made himself look as much as possible like a doctor... (fables1) c. a cat [...] forced himself [to] look as much as possible like a doctor... (4) a. he caused cancer in one woman. (spok: the five 5:00 pm est; 2013) b. *he made cancer [happen] in one woman. c. *he forced cancer [to happen] in one woman. that these verbs appear not to be mutually replaceable is of interest, especially since some prior experimental work has analyzed them as have similar meanings, e.g. that cause, make, and force indicate that the causee did not have a tendency towards the result, the causer and causee were not in concordance, and that the result actually occurred (klettke & wolff 2003, wolff et al. 2005). there is also extensive work arguing that there are important semantic distinctions between these three verbs. for example, nadathur & lauer (2020), lauer & nadathur (2018) argue that make denotes causal sufficiency, while cause denotes causal necessity. furthermore, shibatani (1976) proposes that periphrastic verbs lay on a continuous scale of directness, with varying degrees of control and agency exerted by the causer over the causee. similarly, childers (2016) argues that periphrastic causatives can be ordered on a single causee inclination continuum, in which force denotes the most direct compulsion. this is described as when the causee is non-cooperative, but has no right of refusal. evidently, both the concepts of causee inclination and direct compulsion are related to our previous discussion of sufficiency, since a greater causee inclination requires a smaller degree of compulsion from the causer to bring about the result, and directly compulsing the causee to bring about an intended result is completely sufficient for bringing about the result. notably, much of this aforementioned work makes use of the logics of structural causal models (scms) from pearl (2009), which has previously been used to model causal relations between events as well as their counterfactual values. in our paper, we focus on the constructions x caused/made/forced y (to) z and argue that the relationship between the verbs cause, make, and force is structured not by sufficiency, intentionality, or alternatives alone, but by some interactions of (at least) these three. in order to support our argument, we run an experiment in which participants’ judgements of when the three verbs are appropriate in describing tic-tac-toe sequences is predicted by measures defined using the logics of structural causal models (pearl 2009). 2. possible scales. consider examples (5)–(7). (5) a. i caused martha to go to the gym by mentioning how the habit has helped me. b. *i made martha go to the gym by mentioning how the habit has helped me. c. *i forced martha to go to the gym by mentioning how the habit has helped me. (6) a. i caused martha to go to the gym by criticizing her physical appearance. b. ?i made martha go to the gym by criticizing her physical appearance. c. *i forced martha to go to the gym by criticizing her physical appearance. (7) a. i caused martha to go to the gym by holding her child hostage. b. i made martha go to the gym by holding her child hostage. c. i forced martha to go to the gym by holding her child hostage. 1https://www.gutenberg.org/cache/epub/21/pg21.txt proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 89 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ it appears that the acceptability of the causal verb being used is modulated by several attributes of the causal relata. specifically, the examples give rise to the question – what are the significant differences between the causing events of (5-a) “mentioning how the habit has helped me”, (5-b) “criticizing her physical appearance”, and (5-c) “holding her child hostage”? there are multiple possible analyses – one approach is that (5-c) affects the patient’s space of safe alternatives in a way that (5-a) and (5-b) do not. that is, in the worlds of (5-a) and (5-b), martha can choose not to go to the gym without drastic consequences. in contrast, in the worlds of (5-c), martha’s child might be killed (if the narrator is truthful). relatedly, then, the causing events of (5-a), (5-b), and (5-c) can also be characterized by varying on how sufficient each was in bringing about the effect of martha going to the gym. intuitively, (c) is most sufficient in bringing about the effect. finally, a possible characterization of (5-a)–(5-c) is that the agent of each varies in how intentional they were for the occurrence of the effect. naturally, the agent in (5-c) seems drastically committed to causing martha to go to the gym, moreso than the narrators of (5-a) and (5-b). based on the aforementioned examples and literature, we postulate that these causatives have a semantics built around threshold values on a continuous scale. we consider three measures that are relevant features of causal relationships based on our previous discussion: number of alternatives (alt), intention (int), and sufficiency (suf). 2.1. structural causal models. how can we quantify these concepts? a natural choice is to use games – particularly, the highly constrained, zero-sum game of tic-tac-toe. at every sequence of consecutive moves, the second player had some number of alternatives to the move they ended up taking, the first player intended the second player to make the move it did to some degree, and the first player’s move was, to some degree, sufficient for bringing about the second player’s action. previous work such as hammond et al. (2023) has instantiated examples of games in the framework of structural causal models (scms; termed structural causal games) for the purpose of formalizing agents and their interactions within a grounded, incentivized situation; furthermore, other work such as halpern & kleiman-weiner (2018), nadathur & lauer (2020), lauer & nadathur (2018) and pearl (2019) has defined concepts such as sufficiency and intention within this framework. thus, we can use scms to measure values of alternatives, intention, and sufficiency in tic-tac-toe sequences, which we then use as predictors in participant judgements of when cause, make, and force are accurate in describing the sequences. we define structural causal models in the sense of pearl (2009). scms carve up causal relationships into a discrete set of independent and dependent variables, with defined mechanisms that structurally define variables’ relationships with one another.2 definition 1 (structural causal models). we define a time-indexed causal modelm to consist of: • exogenous variables (u) where each variable xt has an associated set of values it can take on val(xt). exogenous variables have no parents. • endogenous variables (v) where each variable yt has an associated set of values it can take on val(yt) and a timestep t ∈ {0, 1, 2, . . . }. endogenous variables have parents. 2the following definitions are also used by cao et al. (2023) for developing a semantics of causing, enabling, and preventing verbs. proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 90 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ • causal structure (f) represented by arrows running from “parent” variables to “child” variables, which also encode a node’s value based on the value of its parents. we require that all parents immediately precede their children. equivalently, if pt is a parent of c ′t, then t = t′ − 1. the relevant operation of causal models is an intervention, which fixes the value(s) of some variable(s). this action may have downstream changes, but can not affect upstream variables. this term is useful for our later definition of sufficiency. definition 2 (interventions). an intervention i← i is a partial setting i of variables i. a proposition ϕ is true under an intervention, written i← iφ, if φ is true in the model identical tom except the causal mechanisms of i are set to be constant functions mapping to the values in i. thus, a probabilistic scm is a vanilla scm with a probability distribution across exogenous variables (p). in this treatment of introducing probability into a causal model, the uncertainty is “pulled out” (halpern 2016) of the endogenous variables and inserted into exogenous variables, such that the result is a distribution over possible deterministic settings of the model. 2.1.1. a causal model of tic-tac-toe. for our purposes, consider a basic probabilistic causal model that describes the machinations of tic-tac-toe between two agents, player x and player o. we havemttt = (u ,v ,f ,p), where every setting of u delineates possible unfoldings of a tic-tac-toe game based on external factors (e.g., player knowledge) and v = board, where board = {bl t : 0 ≤ t, l ≤ 8} (indicating time and location indices). additionally, at each timestep t, the board-state at t is specified by valuations of all nine assignments of l at t. these valuations are done considering f : ∀z ∈ bl t, z ∈ {x,o, empty}. f delineates the causal mechanisms of tic-tac-toe, namely that (1) x always makes the first move, (2) x and o alternate turns, and (3) the game continues until either (a) a player is able to win by placing three in a row (including horizontally, vertically, and diagonally), or (b) ∀z ∈ b, z ̸= empty. furthermore, p encodes information about ρ – that is, how likely the players choose the highest-utility move. the highest-utility move given a board-state can be calculated using the minimax algorithm, as depicted in figure 1. the minimax algorithm assumes two players – a player that attempts to maximize their possibility of winning, and a player that attempts to minimize the possibility of the former player winning. each possible terminal board state is given a utility score, which we define to be winner × (emptyspace + 1), where winner is −1 if o wins, 0 if it is tie, and +1 if x wins, while emptyspace is the number of empty spaces left on the board at the time of the terminal board state. the latter half of the expression ensures that games that are won earlier are favored. for the sake of assuming an imperfect and probabilistic player, if there is more than one possible transition from t to t+1, the agent takes move argmax(utilityt) or argmin(utilityt) depending on if it is x or o’s play with probability ρ + 1−ρ n where n is the number of empty spaces, and for all other moves take those with probability 1−ρ n . altering the probabilities of possible choices to reflect a uniform distribution might reflect a less skilled player, or comparatively ρ = 1.0 might reflect the plays of a professional player. 2.2. alternatives. firstly, previous work (frankfurt 1969, pereboom 2000) argues that the number of alternative actions available to the causee can distinguish between causal relationships proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 91 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ 0 1 2 3 4 5 6 7 8 utility= maximize t = 5 0 1 2 3 4 5 6 7 8 utility=-2 0 1 2 3 4 5 6 7 8 utility=+3 0 1 2 3 4 5 6 7 8 utility=0 t = 6 0 1 2 3 4 5 6 7 8 minimize 0 1 2 3 4 5 6 7 8 0 1 2 3 4 5 6 7 8 0 1 2 3 4 5 6 7 8 t = 7 utility=+1 utility=0utility=0utility=-2 maximize 0 1 2 3 4 5 6 7 8 0 1 2 3 4 5 6 7 8 0 1 2 3 4 5 6 7 8 terminal (t = 8) utility=0utility=+1utility=0 0.1 0.8 0.1 0.8 0.2 0.80.2 1.0 1.01.0 figure 1: iterating through a partial game using the minimax algorithm, where maximizer= x and minimizer= o. assuming ρ = 0.8, p (x wins) = 0.82 at t = 5. in which the causer is (or is not) culpable for the action taken by the causee. this is also related to lewis (1973)’s argument where actual causation is determined by looking at nearby possible worlds. furthermore, and related to our upcoming discussion of intention, previous work (halpern & kleiman-weiner 2018) has argued that an action taken by an agent when the agent could not do otherwise can never be intentional. this feature is also of interest for differentiating the semantics of causal verbs, since it provides the contrast in (8). (8) a. the child was made to get into the car, although she could’ve chosen to do otherwise. b. ?the child was forced to get into the car, although she could’ve chosen to do otherwise. in tic-tac-toe, a higher number of empty squares signifies a higher degree of freedom and a lesser degree of influence of the causer, while a lower number of empty squares indicates fewer alternatives, suggesting a stronger influence. in this way, the number of potential moves in a tictac-toe board-state can be a rudimentary yet illustrative measure to concretize the continuum of causal influence denoted. so, our first measure alt, which we expect to factor into the predictions of participant judgements of when cause, make, and force are acceptable, quantifies the number of alternative actions available to the causee. our upcoming experiment makes use of tic-tac-toe games as stimuli. in three-state tic-tac-toe sequences as in fig. 2a, alt is measured as the number of alternative actions the agent could have taken, excluding the action that was actually taken. so, alt(y1) = 5. 2.3. intention. secondly, the concept of intention has been argued to be related to alternatives (widerker & mckenna 2003) and also relevant for distinguishing causal situations (copley 2018). proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 92 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ x1 0 1 2 3 4 5 6 7 8 y1 0 1 2 3 4 5 6 7 8 z1 0 1 2 3 4 5 6 7 8 (a) x1 0 1 2 3 4 5 6 7 8 y2 0 1 2 3 4 5 6 7 8 z2 0 1 2 3 4 5 6 7 8 (b) figure 2: examples of three-state tic-tac-toe sequences consider the following sentences: (9) a. john caused the children to dance, but he didn’t intend for the children to dance. b. john made the children dance, but he didn’t intend for the children to dance. c. john forced the children to dance, but he didn’t intend for the children to dance. the intuition is that it is easiest to imagine a situation where (9-a) is true, then perhaps a bit more difficult to imagine a situation where (9-b) is true, and most difficult to imagine a situation where (9-c) is true. we can also see that these sentences are more acceptable with a hedge such as accidentally modifying the main verbs as in (10). (10) a. john accidentally caused the children to dance. he didn’t intend for the children to dance. b. john accidentally made the children dance. he didn’t intend for the children to dance. c. john accidentally forced the children to dance. he didn’t intend for the children to dance. building on this intuition, our second measure (int) is a simplified version of the ‘degree of intention’ proposed by halpern & kleiman-weiner (2018), which is defined within the framework of structural causal models. first, assume a causal modelm, a partial setting of the variables in m, u⃗, an action a⃗, a goal g⃗, and a utility function u(wm, a ← a⃗, u⃗). let u′(wm, a ← a⃗, u⃗) = eu(wm,a←a⃗,u⃗), so that an agent’s expected utility is strictly positive. formally, our definition of int is: int(m, a⃗, g⃗,u′) = pr((m, u⃗) |= (a = a⃗ ∧g = g⃗))u′(wm, a← a⃗, u⃗)∑ (m,u⃗)∈θ:(m,u⃗)|=(a=a⃗′∧g=g⃗) pr(m, u⃗)u′(wm, a← a⃗′, u⃗) in prose, int is the probability that an action performed in a state will result in the desired outcome, normalized by the probability of all alternative actions that would have resulted in the same outcome. first, we scale the utility values by the probability of the desired outcome relative to that action. then, we normalize over all possible utility valuations (for actions for which the goal-state remains a possibility) to consider for comparative cases. this definition captures our intuition that with respect to our tic-tac-toe examples in fig. 2a and fig. 2b, in the statement player o placed at location 2, player o is more intentional in taking this action in z1 than in z2. specifically, any alternative to player o placed at location 2 in fig. 2a, e.g. player o placed at location 5, would make it highly probable that player x wins at the next time-step, thereby largely decreasing the probability of reaching the goal-state of player o. the same is not true for fig. 2b. more generally, our definition takes into consideration that the degree proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 93 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ of intention is higher when the chosen action is more critical to achieving the desired outcome compared to alternatives. since our goal is to use this measure as a predictor measured across tic-tac-toe sequences, where one “goal” (i.e., winning) is as morally good as another, we do not take into consideration side cases involving morality (knobe 2003) that were central to halpern & kleiman-weiner (2018)’s definition. 2.4. sufficiency. thirdly, the notion of causal sufficiency has been well-represented in previous literature on causal verb selection – glass (2023) argues that cause entails local sufficiency, while nadathur & lauer (2020) and lauer & nadathur (2018) argue that make conveys causal sufficiency. as suggested by force-theoretic work on causative verbs (copley & harley 2015, talmy 1988, wolff 2007), this distinguishes between causing and enabling verbs. consider the following examples. (11) a. the pirate made the prisoner walk down the plank. b. the pirate let the prisoner walk down the plank. we can say with certainty that the prisoner walks down the plank in (11-a), while it is less clear whether this result is guaranteed in (11-b). thus, our third model (suf) is pearl (2019)’s “probability of sufficiency”, which is defined as the probability that an event would be sufficient to produce an outcome. descriptively, suf denotes the capacity of c to produce the outcome e in situations where the agent of c did some action other than the one encoded in c. intuitively, player x placing at location 1 in y1 is more sufficient in bringing about player o placing at location 2, than player x placing at location 7 in y2 is for bringing about the same. this is because in sequences where settings y1 and y2 don’t result in player o placing at location 2 at the next time-step, it is more likely that y1 will eventually lead to player o placing at location 2 to block x’s clear three-in-a-row than y2, which does not present that danger to player o. within the framework of scms, pearl (2019, 2009) proposes the probability of sufficiency (suf) mainly to explain why in the case when the presence of oxygen and a lit match are necessary for the occurrence of a fire, it is more felicitous to say that the lit match caused the fire than the oxygen caused the fire. as pearl (2019) writes, the judgement is so because the presence of a lit match is more likely to be sufficient for the fire than the presence of oxygen. assume that we have a causal modelm, a causer action x⃗ ̸∈ u⃗, where u⃗ is a partial setting of the variables inm, and a causee action y⃗. furthermore, u⃗ should have variable assignments for x and y (where x ̸= x⃗). then, suf is defined as: suf(m, x⃗, y⃗, u⃗) = pr(wm,u⃗,y=y⃗ | wm,u⃗,x←x⃗), which is how likely it is for y to become y = y⃗ if x were to counterfactually change from x ̸= x⃗ to x = x⃗. thus, suf quantifies the ability of x = x⃗ to produce the outcome y = y⃗ in situations where y ̸= y⃗. performing this calculation requires a three step process: abduction (updating the prior probabilities in light of m and u⃗), action (intervening on x), and prediction in which we calculate the probability of y = y⃗ given the updated variable values. note that although the definition of suf does not explicitly take into account the utility of an action, our implementation of probability across the tic-tac-toe game-tree does. going back to the examples from fig. 2, observe that “player x forced player o to place at location 2” is more felicitous in fig. 2a than fig. 2b. recall that we predict a higher degree of proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 94 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ sufficiency to be somehow correlated with a higher degree of acceptability of force. assume the notation that b0 . . .b8 refers to each gridcell of a tic-tac-toe state and can take on values from {x,o, empty}. to support our prediction about the acceptability of force, we want to compare the abilities of (1) b1 = x in producing b2 = o and (2) b7 = x in producing b2 = o from fig. 2. with regards to (1), recall that our context u⃗ cannot include b1 = x . so, assume the context is y2 from fig. 2b, which is y2 = (b0 = x)∧ (b6 = o)∧ (b7 = x). however, after conditioning to have b2 ← x , we end up with variable settings equivalent to y1. then, suf(mttt,b1 = x,b2 = o, y2) = 1.0 in the case of ρ = 1.0 (where ρ is the probability that the players choose the highestutility move). it is clear that b2 = o is the highest-utility move in this case because otherwise, player o would lose. next, we compare this with (2), the probability that b7 = x produces b2 = o. we make the context of this y1. however, after conditioning to have b7 = x , we have y2. according to the minimax algorithm (see details in section 2.1), the next best move for playero is now either location 2 or 8. assuming the same ρ, the probability of sufficiency of y2 to produce b2 = o is thus 0.5. since the probability of sufficiency for player x’s move for producing z2 is greater than the probability of sufficiency for the same move to produce z1, this aligns with our intuition that the expression “player x forced player o to place at location 2” is more felicitous in fig. 2a than fig. 2b. there is, however, an exception to this prediction. assume instead that ρ = 0.0, meaning that the players choose their next move at random (e.g. assume that the players are infants). then suf(mttt,b1 = x,b2 = o, y2) = suf(mttt,b7 = x,b2 = o, y2) and the prediction is that neither expression should be preferred. 3. current experiment. in this experiment, we test the three measures described above by creating a dataset of tic-tac-toe games encoding a range of alt, int, and suf values, and evaluating their ability to predict the judgments. the data, analysis scripts, and experimental materials can be accessed here. 3.1. stimuli. in order to generate the 30 stimuli, we first generated 21 full games of tic-tac-toe using a simulated player-and-opponent ran to fulfill a 5x5 design. the first dimension was how many turns it took for the game to complete, measured by the number of empty spaces δ remaining at the end of the game with δ ∈ {1, 2, 3, 4, 5}. the second dimension measured how optimal ρ each player was in selecting the highest-utility play as calculated via the minimax algorithm, with ρ ∈ {0, 0.25, 0.5, 0.75, 1}. this results in 21 games (and not 25) because it is unlikely that a game continues until all squares of a tic-tac-toe board are filled unless ρ = 1. each possible 3-frame sequence from these 21 games were collected in order to create a set of 128 3-step frames. these 128 possible stimuli were then annotated with alt, suf, and int values. since suf and int values skewed towards the lower end of the scale, and since we are primarily interested in suf and int instances towards the higher end of the scale, we define a process to select stimuli where suf and int values differ the most, as well as cover the entire range of suf and int values. we first calculate the absolute difference between the suf and int values for each stimuli. then, we define two sets of ten bins between 0 and 1 and assign each suf and int value into one of these bins. for each bin, we select 3 stimuli with the highest differences in suf and int values. this results in 30 stimuli. in order to artificially grow this set to 60, we proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 95 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ rotate each game board 90◦ clockwise in order to increase the diversity of the target set. 3.2. participants. 109 native english speakers from the us and uk were recruited from prolific. 19 participants were excluded from our analysis for failing at least one of two attentions check that asked whether a specific location was placed with a marker, given an image. out of the remaining 90 participants, age: mean = 36.1, sd = 14.53; gender: 50 female, 38 male, 2 non-binary. each participant was provided with an introduction to the study and had to pass a simple comprehension question about tic-tac-toe to continue. failing the comprehension check brought the participants back to the introductory instructions, after which they could re-attempt the comprehension question. of those that passed the attention checks and comprehension questions, participants took on average 7 minutes and 18 seconds (sd = 3 minutes 42 seconds) to complete the task and were compensated 1.2 gbp. figure 3: example of experiment question 3.3. procedure. participants were first shown a simple explanation of the rules for tic-tac-toe, and then presented with a comprehension question which asked the participant to select the “most likely” next move for a player that must choose a specific location in order to avoid losing. after answering this correctly, participants were presented with 20 pages that had one question each, where one page included both a tic-tac-toe stimulus and a sentence using one of the three causal verbs. the participants were asked to select whether the sentence using cause, make, or force was “accurate” or “inaccurate” in describing the stimulus. an example is shown in fig. 3. of the 20 questions, two were attention checks. the attention checks were designed to appear like the target questions, except participants were asked whether a player placed at a certain location. 4. results & analysis. firstly, we find that holding the set of stimuli constant, participants were less likely to determine made than caused as accurate in describing a scenario, and less likely to determine forced than made as accurate (fig. 5). however, it was not the case that each stronger predicate’s use was a subset of its weaker relatives, and so it is not clear whether these verbs lie proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 96 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ figure 4: proportion of “yes” with 95% cis figure 5: probability of a “yes” rating as a function of each pair of alt, suf, and int in an asymmetric entailment relation (de marneffe et al. 2010). moving on, we first note that the anticorrelation between suf and alt is extremely strong (−0.81), which causes collinearity. this is somewhat expected, since the smaller number of alternatives that the causee has, the more sufficient the causer’s action is for bringing about the result. to ensure stable coefficient estimates, we residualize suf by alt. this means that we take the vertical distance from the line to each of those points as using that distance as the predictor, rather than including the probability of sufficiency itself. the resulting predictor (sufresidalt) is interpreted as capturing all the information that suf provides that is not shared with alt. we fit multiple bayesian regressions with a bernoulli family, using participant judgements as the outcome variable. in our first model (i; full model results in table 2), we include a full three-way interaction between the value of verb, int, alt, and sufresidalt, as well as random effects for verb and participant. the random effects accounts for variability at the verb and participant levels. in our second model (ii), we include all two-way interactions but exclude the three-way interaction between the continuous variables, but otherwise keep the random effects present in (i). in our third model (iii), we do not include interactions between verb and the continuous variables, but keep the three-way interaction among the continuous variables (sufresidalt, int, suf) as well as the categorical verb predictor and the random effects present in (i). our last model (iv) includes only two-way interactions among the continuous variables and excludes the three-way interaction, but keeps the categorical verb predictor and the random effects present in (i). we then compare these four models. we first find that the difference in elpd (expected log proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 97 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ predictive density) between model iii and model iv is −20.9, with a se of 6.6. since iii has a higher elpd, it is the better model. the large negative difference suggests that dropping the three-way interaction among the continuous variables (as in iv) significantly reduces the model’s predictive accuracy. this means that the interactions between sufresidalt, int, and suf provides important information for optimizing predictions of when participants describe uses of cause, make, and force as accurate. next, the difference in elpd between i and ii is −21.8, with a se of 7.1. i has a higher elpd, indicating that it is the better model. again, removing the threeway interaction among the continuous variables (in ii) significantly reduces the model’s predictive performance. finally, we compare models i and iii. the difference in elpd between iii and i is −3.1, with a se of 4.8. the difference is relatively small and the se is larger than the absolute difference, meaning this difference is not significant. therefore, both models perform similarly, with no clear advantage of including the interactions between verb and the continuous variables as in the full model (i). firstly, every standard deviation increase in int causes the log-odds of a “yes” for cause to go up by ∼ 0.5. this means that a stimuli with a higher degree of int is more likely to have cause, make, and force rated as accurate. also, every standard deviation increase in residualized suf predictor causes the log-odds of a “yes” for cause to go up by∼ 1.2, which is more than twice the size of the effect of int. so, a higher degree of residualized suf increases the probability of accuracy much more than increasing int does. next, every standard deviation increase in alt causes the log-odds of a “yes” for cause to go down by ∼ 0.8. it is expected that this effect is in the opposite direction than the other two predictors, since alt is anticorrelated with suf. we also note that when alt matches in sign with residualized suf or int, the log-odds of a “yes” for cause goes up ∼ 0.5 in the produce of standard deviations. interestingly, there is a comparatively large and reliable interaction between the made level of verb with residualized suf and int, unlike the other levels of verb. the same observation holds for the verb cause, and the interaction between sufresidalt and alt, as well as int and alt (see the highlighted rows in table 2). these interactions suggest that made has some additional semantic component that is also a function of suf and int, that is not present in the other verbs. the same possibility holds for cause and its notable interactions. in order to explore this in more detail, we fit three additional models for each verb. after subsetting the data by verb, model v predicts participants’ judgments of stimuli that use the verb caused by including three continuous predictors (as fixed effects), along with their interactions. the model also includes random intercepts for participant to account for variability in individual responses. model vi and vii are similar, except vi only pertains to stimuli that use made, and vii only pertains to stimuli that use forced (see all complete model results in tabs. 3–5). as we expect, all three models reliably use our three continuous measures as predictors of participant judgements of each verb. more interestingly, each verb has a unique combination of interactions that are reliable (see tab. 1). in model v, the interactions of continuous predictors that are reliable for cause are firstly, the interaction between sufresidalt and alt; secondly, the interaction between int and alt; and thirdly, the three-way interaction between sufresidalt, int, and alt. as for made in model vi, surprisingly, the interactions that are reliable are not reflected in the original full model (i). the reliable predictors for judgements of made include the proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 98 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ caused made forced sufresidalt:int + sufresidalt:alt + + + int:alt + + + sufresidalt:int:alt table 1: estimates of interactions’ intercepts by verb. light grey indicates an unreliable effect. interaction between sufresidalt and alt, as well as sufresidalt, int, and alt. finally, in model vii, the interaction between sufresidalt, int, and alt are reliable. 5. discussion. to begin, our results support the prediction that intention, sufficiency, and possible alternative actions factor into the semantics of the causal senses of cause and force, which is demonstrated by the reliable intercepts for these level of verb in model i. furthermore, each continuous measure shows as reliable in this model. it is less clear whether this holds for made since its credible interval includes 0 in the full model (i). however, we observe that the smaller madespecific model (vi) reliably uses all three of sufresidalt, int, and alt as predictors. the uncertainty in model i may originate from a variety of factors – for example, even interactions that are reliable for made, such as its interaction with sufresidalt and int, have wide credible intervals (0.02 to 0.87), which may contribute to higher overall uncertainty. although the results for made are unclear, it seems unlikely that the concepts do not at all contribute to the semantics of made. regarding alternatives, e.g., consider one participant of our task’s post-survey comment: “if there were multiple options, [i.e.] more than one blank spot where the player could select, [...] i made an assumption that “forced” or “made” were inaccurate.” evidently, participants take into consideration the number of alternatives that a causee has when judging the accuracy of make. taking a wider view, that each verb takes a unique combination of interactions supports the argument that these components convey distinct information to our models of participant judgements. furthermore, the semantics of each verb takes into consideration a unique blend of each concept. it is also interesting to observe that cause, make, and force have a decreasing number of reliable interactions that are used as predictors. speculatively, these results seem to convey that what we perceive as increasing “force” in these verbs is actually a conglomerate of multiple factors. to conclude, our work has provided experimental support for our hypothesis that causal verbs such as cause, make, and force exhibit distinct and nuanced semantics, which are shaped by a combination of factors including sufficiency, intention, and alternatives. through our experimental analysis, we demonstrated that no single predictor, such as sufficiency alone, fully determines the appropriateness of these verbs. instead, the interactions between these factors play a unique role in determining how participants judge the accuracy of each of these verbs in describing various causal scenarios. proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 99 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ references baglini, rebekah & elitzur a bar-asher siegal. 2021. modelling linguistic causation. manuscript, aarhus university and hebrew university of jerusalem . cao, angela, atticus geiger, elisa kreiss, thomas icard & tobias gerstenberg. 2023. a semantics for causing, enabling, and preventing verbs using structural causal models. in proceedings of the 45th annual conference of the cognitive science society, 2947–2954. cao, angela, gregor williamson & jinho choi. 2022. a cognitive approach to annotating causal constructions in a cross-genre corpus. in proceedings of the 16th linguistic annotation workshop (law) at lrec, 151–159. online: european language resources association. http://lrec-conf.org/proceedings/lrec2022/workshops/lawxvi/pdf/2022.lawxvi-1.18.pdf. childers, zachary. 2016. cause and affect: evaluative and emotive parameters of meaning among the periphrastic causative verbs in english. austin, tx: university of texas, austin phd dissertation. copley, bridget. 2018. dispositional causation. glossa: a journal of general linguistics 3. 10.5334/gjgl.507. copley, bridget & heidi harley. 2015. a force-theoretic framework for event structure. linguistics and philosophy 38(2). 103–158. 10.1007/s10988-015-9168-x. davies, mark. 2008–. the corpus of contemporary american english (coca). available online at https://www.english-corpora.org/coca/. frankfurt, harry g. 1969. alternate possibilities and moral responsibility. the journal of philosophy 66(23). 829–839. http://www.jstor.org/stable/2023833. glass, lelia. 2023. using the anna karenina principle to explain why cause favors negativesentiment complements. semantics and pragmatics 16(6). 1–48. 10.3765/sp.16.6. halpern, joseph y. 2016. actual causality. the mit press. https://doi.org/10.7551/mitpress/ 10809.001.0001. halpern, joseph y. & max kleiman-weiner. 2018. towards formal definitions of blameworthiness, intention, and moral responsibility. hammond, lewis, james fox, tom everitt, ryan carey, alessandro abate & michael wooldridge. 2023. reasoning about causality in games. artificial intelligence 320. 103919. klettke, bianca & philip wolff. 2003. differences in how english and german speakers talk and reason about cause. in proceedings of the annual meeting of the cognitive science society, vol. 25, 675–680. knobe, joshua. 2003. intentional action and side effects in ordinary language. analysis 63. 190– 193. lauer, sven & prerna nadathur. 2018. sufficiency causatives. unpublished manuscript. lewis, david. 1973. causation. journal of philosophy 70(17). 556–567. 10.2307/2025310. de marneffe, marie-catherine, christopher d. manning & christopher potts. 2010. “was it good? it was provocative.” learning the meaning of scalar adjectives. in jan hajič, sandra carberry, stephen clark & joakim nivre (eds.), proceedings of the 48th annual meeting of the association for computational linguistics, 167–176. uppsala, sweden: association for computational linguistics. https://aclanthology.org/p10-1018. nadathur, prerna & sven lauer. 2020. causal necessity, causal sufficiency, and the implications proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 100 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ of causative verbs. glossa: a journal of general linguistics 5. 49–105. nadathur, prerna & elitzur a bar-asher siegal. 2022. modeling progress: causal models, event types, and the imperfective paradox. in west coast conference in formal linguistics (wccfl), vol. 40, . pearl, judea. 2009. causality: models, reasoning and inference. cambridge university press 2nd edn. pearl, judea. 2019. sufficient causes: on oxygen, matches, and fires. journal of causal inference 7(2). 20190026. https://doi.org/10.1515/jci-2019-0026. pereboom, derk. 2000. alternative possibilities and causal histories. philosophical perspectives 14. 119–137. http://www.jstor.org/stable/2676125. schulz, katrin. 2011. if you’d wiggled a, then b would’ve changed: causality and counterfactual conditionals. synthese 179(2). 239–251. 10.1007/s11229-010-9780-9. shibatani, m. 1976. the grammar of causative constructions syntax and semantics online, isbn: 9789004425774. brill. https://books.google.com/books?id=ft2kzgeacaaj. talmy, leonard. 1988. force dynamics in language and cognition. cognitive science 12(1). 49– 100. https://doi.org/10.1016/0364-0213(88)90008-0. https://www.sciencedirect.com/science/ article/pii/0364021388900080. widerker, david & michael mckenna (eds.). 2003. moral responsibility and alternative possibilities: essays on the importance of alternative possibilities. ashgate. williamson, gregor, angela cao, yingying chen, yuxin ji, liyan xu & jinho d. choi. 2023. exploring a multi-layered cross-genre corpus of document-level semantic relations. information 14(8). 431. 10.3390/info14080431. wolff, phillip. 2007. representing causation. journal of experimental psychology. general 136. 82–111. 10.1037/0096-3445.136.1.82. wolff, phillip, bianca klettke, tracy ventura & geunyoung song. 2005. expressing causation in english and other languages. in woo-kyoung ahn, robert l. goldstone, bradley c. love, arthur b. markman & phillip wolff (eds.), categorization inside and outside the laboratory: essays in honor of douglas l. medin, 29–48. american psychological association. 10.1037/11156-003. proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 101 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ appendix parameter estimate est. error l-95% ci u-95% ci intercept -0.43 0.15 -0.74 -0.14 verbforced -0.67 0.21 -1.08 -0.26 verbmade -0.25 0.19 -0.61 0.11 sufresidalt 1.19 0.16 0.89 1.50 int 0.54 0.13 0.28 0.81 alt -0.82 0.14 -1.11 -0.55 sufresidalt:int -0.30 0.16 -0.61 0.02 sufresidalt:alt 0.44 0.16 0.13 0.76 int:alt 0.50 0.18 0.16 0.86 verbforced:sufresidalt 0.09 0.23 -0.36 0.55 verbmade:sufresidalt -0.32 0.22 -0.74 0.10 verbforced:int 0.11 0.20 -0.28 0.51 verbmade:int -0.11 0.19 -0.47 0.25 verbforced:alt 0.10 0.21 -0.34 0.52 verbmade:alt 0.05 0.19 -0.33 0.44 sufresidalt:int:alt -0.72 0.19 -1.07 -0.35 verbforced:sufresidalt:int 0.32 0.23 -0.15 0.77 verbmade:sufresidalt:int 0.45 0.22 0.02 0.87 verbforced:sufresidalt:alt -0.34 0.24 -0.80 0.13 verbmade:sufresidalt:alt -0.03 0.22 -0.46 0.40 verbforced:int:alt -0.40 0.30 -1.00 0.19 verbmade:int:alt -0.39 0.25 -0.87 0.11 verbforced:sufresidalt:int:alt -0.42 0.29 -0.99 0.15 verbmade:sufresidalt:int:alt 0.26 0.26 -0.24 0.76 table 2: parameter estimates for model i. proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 102 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/ parameter estimate est. error l-95% ci u-95% ci intercept -0.46 0.15 -0.76 -0.17 sufresidalt 1.21 0.17 0.90 1.54 int 0.54 0.14 0.27 0.82 alt -0.83 0.14 -1.13 -0.56 sufresidalt:int -0.28 0.16 -0.60 0.04 sufresidalt:alt 0.46 0.17 0.14 0.79 int:alt 0.52 0.18 0.16 0.89 sufresidalt:int:alt -0.71 0.19 -1.07 -0.35 table 3: parameter estimates for model v. parameter estimate est. error l-95% ci u-95% ci intercept -0.65 0.14 -0.93 -0.39 sufresidalt 0.87 0.15 0.58 1.17 int 0.44 0.13 0.19 0.69 alt -0.74 0.13 -1.00 -0.50 sufresidalt:int 0.16 0.15 -0.13 0.45 sufresidalt:alt 0.41 0.15 0.12 0.70 int:alt 0.09 0.17 -0.23 0.43 sufresidalt:int:alt -0.44 0.17 -0.78 -0.10 table 4: parameter estimates for model vi. parameter estimate est. error l-95% ci u-95% ci intercept -1.13 0.17 -1.48 -0.80 sufresidalt 1.26 0.18 0.92 1.61 int 0.63 0.15 0.35 0.92 alt -0.70 0.17 -1.04 -0.38 sufresidalt:int -0.01 0.17 -0.34 0.32 sufresidalt:alt 0.08 0.18 -0.27 0.42 int:alt 0.12 0.24 -0.35 0.59 sufresidalt:int:alt -1.15 0.23 -1.61 -0.71 table 5: parameter estimates for model vii. proceedings of elm 3: 88-103, 2025 angela cao, aaron steven white, and daniel lassiter: cause, make, and force as graded causatives. 103 https://journals.linguisticsociety.org/proceedings/index.php/elm/issue/archive https://www.elm-conference.net/